about
[Re:DeepSeek-OCR] Optical Context Compression Is Just (Bad) Autoencoding (arxiv.org)
6 points by atbhtunnm 298 days ago | hide | past | pdf | 1 comment on HN

In plain words: They tested whether turning a model's stored text representations into pictures and squeezing them with a vision system beats simpler ways to shrink text. It does not: plain averaging or a learned encoder matched or beat it, and it only tied with dropping text.

Abstract · Optical Context Compression Is Just (Bad) Autoencoding

DeepSeek-OCR shows that rendered text can be reconstructed from a small number of vision tokens, sparking excitement about using vision as a compression medium for long textual contexts. But this pipeline requires rendering token embeddings to pixels and compressing from there -- discarding learned representations in favor of an image the vision encoder must then recover from. We ask whether this detour helps. Comparing DeepSeek-OCR's vision encoder against near-zero-parameter mean pooling and a learned hierarchical encoder, we find it does not. For reconstruction, simple direct methods match or surpass vision at every compression ratio. For language modeling, vision performs comparably to truncation -- a baseline that simply discards context -- and loses to the hierarchical encoder at every compression ratio. As expected, all compression methods outperform truncation for factual recall, but vision never surpasses the best direct baseline. The excitement around optical context compression outpaces the evidence. Code and checkpoints are available at https://github.com/ivnle/bad-autoencoding.

Ivan Yee Lee, Cheng Yang, Taylor Berg-Kirkpatrick
arXiv:2512.03643 · cs.CV, cs.CL, cs.LG · submitted Dec 3, 2025 · updated Apr 4, 2026
abstract · pdf · html

add comment on HN
Also discussed: Dec 2025 (21 points, 1 comment)

Author here. I started this project after reading the earlier threads on DeepSeek-OCR [1][2]. I got really excited about "vision for context compression," but after reading their paper, a couple things were bugging me.

They show good OCR results (image → text), but the pitch is context compression (text → image → text). They never actually test that pipeline. So I implemented it: render text, compress to vision tokens, reconstruct. Then I compared against just compressing the text embeddings directly. Mean pooling (averaging embeddings in a sliding window) nearly matched DeepSeek-OCR. A small conv encoder crushed both.

Ok so fine, maybe vision isn't special for reconstruction. But maybe the path matters more than the destination. Do the representations learned through vision work better for language modeling? I finetuned the checkpoints from the reconstruction experiments for next-token prediction. Vision and mean pooling couldn't beat truncation, but the conv encoder could. I didn't do any architecture search. It just worked.

That said, this is preliminary work. I just wanted to answer the obvious next questions. So far, the findings don't support the "vision for context compression" narrative.

Happy to answer questions.

[1] https://news.ycombinator.com/item?id=45640594 [2] https://news.ycombinator.com/item?id=45658928