about
Conflict Adaptation in Vision-Language Models (arxiv.org)
1 point by PaulHoule 314 days ago | hide | past | pdf | discuss on HN

In plain words: When a word's meaning clashes with its ink color, vision-language models did better on a hard trial after another hard one, a human sign of adjusting control. 12 of 13 showed this, and removing one model's inner unit raised errors only on clashing trials.

Abstract

A signature of human cognitive control is conflict adaptation: improved performance on a high-conflict trial following another high-conflict trial. This phenomenon offers an account for how cognitive control, a scarce resource, is recruited. Using a sequential Stroop task, we find that 12 of 13 vision-language models (VLMs) tested exhibit behavior consistent with conflict adaptation, with the lone exception likely reflecting a ceiling effect. To understand the representational basis of this behavior, we use sparse autoencoders (SAEs) to identify task-relevant supernodes in InternVL 3.5 4B. Partially overlapping supernodes emerge for text and color in both early and late layers, and their relative sizes mirror the automaticity asymmetry between reading and color naming in humans. We further isolate a conflict-modulated supernode in layers 24-25 whose ablation significantly increases Stroop errors while minimally affecting congruent trials.

Xiaoyang Hu
arXiv:2510.24804 · cs.CV, cs.CL · submitted Oct 28, 2025 · updated Nov 18, 2025
abstract · pdf · html · Workshop on Interpreting Cognition in Deep Learning Models at NeurIPS 2025

add comment on HN