about
Conditional Hallucinations for Image Compression (arxiv.org)
1 point by Hard_Space on Oct 30, 2024 | hide | past | pdf | discuss on HN

In plain words: A compression model decides how much invented detail to add for each image, using a learned guess of what people prefer to tune its training loss. It beat today's best image compressors, which pick one fixed level of invention for every picture.

Abstract

In lossy image compression, models face the challenge of either hallucinating details or generating out-of-distribution samples due to the information bottleneck. This implies that at times, introducing hallucinations is necessary to generate in-distribution samples. The optimal level of hallucination varies depending on image content, as humans are sensitive to small changes that alter the semantic meaning. We propose a novel compression method that dynamically balances the degree of hallucination based on content. We collect data and train a model to predict user preferences on hallucinations. By using this prediction to adjust the perceptual weight in the reconstruction loss, we develop a Conditionally Hallucinating compression model (ConHa) that outperforms state-of-the-art image compression methods. Code and images are available at https://polybox.ethz.ch/index.php/s/owS1k5JYs4KD4TA.

Till Aczel, Roger Wattenhofer
arXiv:2410.19493 · eess.IV, cs.CV, cs.LG · submitted Oct 25, 2024 · updated Mar 5, 2025
abstract · pdf · html

add comment on HN