about
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition (arxiv.org)
2 points by PaulHoule on Jul 8, 2025 | hide | past | pdf | discuss on HN

In plain words: They tested vision-language emotion readers on face photos labeled for whether teeth show, to check which visual cues drive their guesses. Accuracy shifted with visible teeth, and the best model leaned on eyebrow position, showing it uses shortcuts rather than grounded emotion cues.

Abstract

Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual cues do these models rely on to infer affect, and are these cues psychologically grounded or superficially learnt? We benchmark varying scale VLMs on a teeth annotated subset of AffectNet dataset and find consistent performance shifts depending on the presence of visible teeth. Through structured introspection of, the best-performing model, i.e., GPT-4o, we show that facial attributes like eyebrow position drive much of its affective reasoning, revealing a high degree of internal consistency in its valence-arousal predictions. These patterns highlight the emergent nature of FMs behaviour, but also reveal risks: shortcut learning, bias, and fairness issues especially in sensitive domains like mental health and education.

Iosif Tsangko, Andreas Triantafyllopoulos, Adem Abdelmoula, Adria Mallol-Ragolta, Bjoern W. Schuller
arXiv:2506.19079 · cs.CV, cs.AI, cs.HC · submitted Jun 23, 2025
abstract · pdf · html

add comment on HN