In plain words: They break a language model's tangled internal signals into separate, single-meaning features to see what it understands. When the model seems to misread a prompt, the system rewrites it with clarifying notes, beating the usual ask-it-directly approach on math problems and metaphor spotting.
Abstract
Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection.
Shun Wang, Tyler Loakman, Youbo Lei, Yi Liu, Bohao Yang, Yuting Zhao, Dong Yang, Chenghua Lin
arXiv:2507.06427 · cs.CL, cs.LG · submitted Jul 8, 2025
abstract · pdf · html