about
Optical context compression is just (bad) autoencoding (arxiv.org)
21 points by unclefuzzy 299 days ago | hide | past | pdf | 1 comment on HN

In plain words: They tested whether turning a model's stored text representations into pictures and squeezing them with a vision system beats simpler ways to shrink text. It does not: plain averaging or a learned encoder matched or beat it, and it only tied with dropping text.

Abstract · Optical Context Compression Is Just (Bad) Autoencoding

DeepSeek-OCR shows that rendered text can be reconstructed from a small number of vision tokens, sparking excitement about using vision as a compression medium for long textual contexts. But this pipeline requires rendering token embeddings to pixels and compressing from there -- discarding learned representations in favor of an image the vision encoder must then recover from. We ask whether this detour helps. Comparing DeepSeek-OCR's vision encoder against near-zero-parameter mean pooling and a learned hierarchical encoder, we find it does not. For reconstruction, simple direct methods match or surpass vision at every compression ratio. For language modeling, vision performs comparably to truncation -- a baseline that simply discards context -- and loses to the hierarchical encoder at every compression ratio. As expected, all compression methods outperform truncation for factual recall, but vision never surpasses the best direct baseline. The excitement around optical context compression outpaces the evidence. Code and checkpoints are available at https://github.com/ivnle/bad-autoencoding.

Ivan Yee Lee, Cheng Yang, Taylor Berg-Kirkpatrick
arXiv:2512.03643 · cs.CV, cs.CL, cs.LG · submitted Dec 3, 2025 · updated Apr 4, 2026
abstract · pdf · html

add comment on HN
Also discussed: Dec 2025 (6 points, 1 comment)

Academic papers should be neutral when dissing previous work imo.

"[Previous work] is just bad [previous other work]" isn't a professional way to talk about the merits and drawbacks of competing approaches, but boy does it get tiktok views!

Better work stands on its own merits. No need to explicitly shit on the competition in the title/abstract.