about
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging (arxiv.org)
1 point by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss on HN

In plain words: The system turns brain scans into a detailed text description first, then uses that description to guide an image generator toward what the person saw. Adding this text step recovers more detail and fewer meaning errors than rebuilding the image straight from brain activity.

Abstract

Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic inconsistencies, which may be attributed to insufficient semantic information. To address this issue, we propose an approach named Fine-grained Brain-to-Image reconstruction (FgB2I), which employs fine-grained text as bridge to improve image reconstruction. FgB2I comprises three key stages: detail enhancement, decoding fine-grained text descriptions, and text-bridged brain-to-image reconstruction. In the detail-enhancement stage, we leverage large vision-language models to generate fine-grained captions for visual stimuli and experimentally validate its importance. We propose three reward metrics (object accuracy, text-image semantic similarity, and image-image semantic similarity) to guide the language model in decoding fine-grained text descriptions from fMRI signals. The fine-grained text descriptions can be integrated into existing reconstruction methods to achieve fine-grained Brain-to-Image reconstruction.

Runze Xia, Shuo Feng, Renzhi Wang, Congchi Yin, Xuyun Wen, Piji Li
arXiv:2505.22150 · cs.CV, cs.CL · submitted May 28, 2025 · updated May 29, 2025
abstract · pdf · html · CogSci2025

add comment on HN