about
Compressed-Language Models for Understanding Compressed File Formats: JPEG (arxiv.org)
3 points by jasondavies on May 31, 2024 | hide | past | pdf | discuss on HN

In plain words: Language models were fed the raw bytes of JPEG files—no decoding—and tested on reading file properties, spotting broken files, and writing new ones. They handled all three, showing they can make sense of compressed data without ever unpacking it.

Abstract · Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration

This study investigates whether Compressed-Language Models (CLMs), i.e. language models operating on raw byte streams from Compressed File Formats~(CFFs), can understand files compressed by CFFs. We focus on the JPEG format as a representative CFF, given its commonality and its representativeness of key concepts in compression, such as entropy coding and run-length encoding. We test if CLMs understand the JPEG format by probing their capabilities to perform along three axes: recognition of inherent file properties, handling of files with anomalies, and generation of new files. Our findings demonstrate that CLMs can effectively perform these tasks. These results suggest that CLMs can understand the semantics of compressed data when directly operating on the byte streams of files produced by CFFs. The possibility to directly operate on raw compressed files offers the promise to leverage some of their remarkable characteristics, such as their ubiquity, compactness, multi-modality and segment-nature.

Juan C. Pérez, Alejandro Pardo, Mattia Soldan, Hani Itani, Juan Leon-Alcazar, Bernard Ghanem
arXiv:2405.17146 · cs.CV · submitted May 27, 2024
abstract · pdf · html

add comment on HN