about
How far are we from scaling up next-pixel prediction for image pretraining? (arxiv.org)
2 points by PaulHoule 297 days ago | hide | past | pdf | discuss on HN

In plain words: They trained Transformers to build images pixel by pixel, testing how to divide compute between models and more data. The best split depends on the task: generation needs data growing three to five times faster than classification, and compute, not data, is the limit.

Abstract · Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?

This paper investigates the scaling properties of autoregressive next-pixel prediction, a simple, end-to-end yet under-explored framework for unified vision models. Starting with images at resolutions of 32x32, we train a family of Transformers using IsoFlops profiles across compute budgets up to 7e19 FLOPs and evaluate three distinct target metrics: next-pixel prediction objective, ImageNet classification accuracy, and generation-based completion measured by Fr'echet Distance. First, optimal scaling strategy is critically task-dependent. At a fixed resolution of 32x32 alone, the optimal scaling properties for image classification and image generation diverge, where generation optimal setup requires the data size grow three to five times faster than for the classification optimal setup. Second, as image resolution increases, the optimal scaling strategy indicates that the model size must grow much faster than data size. Surprisingly, by projecting our findings, we discover that the primary bottleneck is compute rather than the amount of training data. As compute continues to grow four to five times annually, we forecast the feasibility of pixel-by-pixel modeling of images within the next five years.

Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu, Yifeng Lu, Quoc V. Le
arXiv:2511.08704 · cs.CV, cs.LG · submitted Nov 11, 2025 · updated May 17, 2026
abstract · pdf · html · Accepted by ICML2026

add comment on HN
Also discussed: Nov 2025 (1 point, 0 comments)