about
High-Quality Video Reconstruction from Brain Activity (arxiv.org)
18 points by kevinventullo on May 24, 2023 | hide | past | pdf | 3 comments on HN

In plain words: A system turns brain activity recorded while people watch videos into moving footage, learning from paired scans and clips to match what the brain sees. Its clips matched the real scenes' meaning 85% of the time, beating the previous best brain-to-video approach by 45%.

Abstract · Cinematic Mindscapes: High-quality Video Reconstruction from Brain Activity

Reconstructing human vision from brain activities has been an appealing task that helps to understand our cognitive process. Even though recent research has seen great success in reconstructing static images from non-invasive brain recordings, work on recovering continuous visual experiences in the form of videos is limited. In this work, we propose Mind-Video that learns spatiotemporal information from continuous fMRI data of the cerebral cortex progressively through masked brain modeling, multimodal contrastive learning with spatiotemporal attention, and co-training with an augmented Stable Diffusion model that incorporates network temporal inflation. We show that high-quality videos of arbitrary frame rates can be reconstructed with Mind-Video using adversarial guidance. The recovered videos were evaluated with various semantic and pixel-level metrics. We achieved an average accuracy of 85% in semantic classification tasks and 0.19 in structural similarity index (SSIM), outperforming the previous state-of-the-art by 45%. We also show that our model is biologically plausible and interpretable, reflecting established physiological processes.

Zijiao Chen, Jiaxin Qing, Juan Helen Zhou
arXiv:2305.11675 · cs.CV, cs.CE · submitted May 19, 2023
abstract · pdf · html · 15 pages, 11 figures, submitted to anonymous conference

add comment on HN
Also discussed: May 2023 (2 points, 0 comments) · May 2023 (2 points, 0 comments) · May 2023 (1 point, 1 comment) · May 2023 (1 point, 1 comment)

Not gonna lie, it's impressive. But that said, courtroom drama this is classic recovered-memory stuff:

"your recovered memory says you saw TWO people walking. Police footage shows ONE person walking"

"case closed: its a false memory"

So its prompted adversarial semantic match is good: birb matches birb. jogger(s) match jogger. cloud matches cloud. But its not faithful image recovery at any stretch. (nor do they claim it. It's how the downstream consumption of this work will what-if on it)

Some of the goodness looks like pure and simple motion recovery: if you put cat == cat and then it's picked up tracking, making it tracking-cat isn't exactly hard.

This sounds and looks impressive but is this actually anything more than a brain machine interface for a stable diffusion like model using an MRI?

If so what it generates and reconstructs is based on the data the model was trained upon not what the individual has actually seen. So this only provides a thematic prompt at best and the rest is just a hallucination.

There is no extraction of actual visual memories from the individual.

I really hope that this won’t going to go any further it can be 1000 times worse than bite analysis if this ever spills into forensic science.

This is wild. I didn't know anyone was seriously working on this let alone it being this far along.