about
Microsoft Paper: LLMs Corrupt Your Documents When You Delegate (Arxiv.org) (arxiv.org)
7 points by wuschel 159 days ago | hide | past | pdf | 2 comments on HN

In plain words: Built a test that simulates long editing jobs handed to AI across 52 professional fields, checking whether documents stay intact after repeated edits. Even the strongest models damaged about a quarter of the content by the end, and extra tools did not help.

Abstract · LLMs Corrupt Your Documents When You Delegate

Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely. Additional experiments reveal that agentic tool use does not improve performance on DELEGATE-52, and that degradation severity is exacerbated by document size, length of interaction, or presence of distractor files. Our analysis shows that current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction.

Philippe Laban, Tobias Schnabel, Jennifer Neville
arXiv:2604.15597 · cs.CL, cs.HC · submitted Apr 17, 2026
abstract · pdf · html

add comment on HN
Also discussed: May 2026 (479 points, 201 comments) · Apr 2026 (4 points, 0 comments) · Apr 2026 (2 points, 0 comments) · Apr 2026 (4 points, 2 comments)

even frontier models (...) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely

Wow, 25% corrupted seems like a lot. The abstract and the intro of this paper emphasizes "documents" and it's Microsoft, so I assumed Word docs, but that's not true, they used a wide variety of things, graphs, text files, possibly images, or some machine readable description of textile weaving. A proof reader might not catch 25% corrupted textile description file, or 25% corruption in a graph.

Is this "corruption" what in text files we've all been taught to call "hallucinations"?

Our analysis shows that current LLMs are unreliable delegates:

Who knew that a tool that relies on probability could make such a mess?