In plain words: A model trained on brain scans from several people learns a shared way to line up their signals, then needs just one hour of a new person's scans to rebuild the images they saw. It beat models trained separately on dozens of hours per person.
Abstract · MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data
Reconstructions of visual perception from brain activity have improved tremendously, but the practical utility of such methods has been limited. This is because such models are trained independently per subject where each subject requires dozens of hours of expensive fMRI training data to attain high-quality results. The present work showcases high-quality reconstructions using only 1 hour of fMRI training data. We pretrain our model across 7 subjects and then fine-tune on minimal data from a new subject. Our novel functional alignment procedure linearly maps all brain data to a shared-subject latent space, followed by a shared non-linear mapping to CLIP image space. We then map from CLIP space to pixel space by fine-tuning Stable Diffusion XL to accept CLIP latents as inputs instead of text. This approach improves out-of-subject generalization with limited training data and also attains state-of-the-art image retrieval and reconstruction metrics compared to single-subject approaches. MindEye2 demonstrates how accurate reconstructions of perception are possible from a single visit to the MRI facility. All code is available on GitHub.
Paul S. Scotti, Mihir Tripathy, Cesar Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A. Norman, Tanishq Mathew Abraham
arXiv:2403.11207 · cs.CV, cs.AI, q-bio.NC · submitted Mar 17, 2024 · updated Jun 15, 2024
abstract · pdf · html · In Forty-first International Conference on Machine Learning, 2024. Code at https://github.com/MedARC-AI/MindEyeV2. Published as a conference paper at ICML 2024
> The present work demonstrates that it is now practical for patients to undergo a single MRI scanning session and produce enough data to perform high-quality reconstructions of their visual perception.
> Such image reconstructions from brain activity are expected to be systematically distorted due to factors including mental state, neurological conditions, etc.
> This could potentially enable novel clinical diagnosis and assessment approaches, including applications for improved locked-in (pseudocoma) patient communication (Monti et al., 2010) and brain-computer interfaces if adapted to real-time analysis (Wallace et al., 2022) or non-fMRI neu-roimaging modalities.
> As technology continues to improve, we note it is important that brain data be carefully protected and companies collecting such data be transparent with their use.