about
Understanding Hidden Computations in Chain-of-Thought Reasoning (arxiv.org)
1 point by riemann77 on Dec 12, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Models can solve hard reasoning problems even when their written steps are replaced by filler dots. Peeking at the layers inside recovers those hidden characters with no loss in accuracy, where reading the visible steps shows nothing.

Abstract

Chain-of-Thought (CoT) prompting has significantly enhanced the reasoning abilities of large language models. However, recent studies have shown that models can still perform complex reasoning tasks even when the CoT is replaced with filler(hidden) characters (e.g., "..."), leaving open questions about how models internally process and represent reasoning steps. In this paper, we investigate methods to decode these hidden characters in transformer models trained with filler CoT sequences. By analyzing layer-wise representations using the logit lens method and examining token rankings, we demonstrate that the hidden characters can be recovered without loss of performance. Our findings provide insights into the internal mechanisms of transformer models and open avenues for improving interpretability and transparency in language model reasoning.

Aryasomayajula Ram Bharadwaj
arXiv:2412.04537 · cs.CL, cs.LG · submitted Dec 5, 2024
abstract · pdf · html

add comment on HN

     Chain-of-Thought (CoT) prompting has significantly enhanced the reasoning abilities of large language models. However, recent studies have shown that models can still perform complex reasoning tasks even when the CoT is replaced with filler(hidden) characters (e.g., "..."), leaving open questions about how models internally process and represent reasoning steps. In this paper, we investigate methods to decode these hidden characters in transformer models trained with filler CoT sequences. By analyzing layer-wise representations using the logit lens method and examining token rankings, we demonstrate that the hidden characters can be recovered without loss of performance. Our findings provide insights into the internal mechanisms of transformer models and open avenues for improving interpretability and transparency in language model reasoning.