about
ThirdEye: Brain-Inspired Mono Depth-Estimation (arxiv.org)
4 points by teocalin37 on Jun 27, 2025 | hide | past | pdf | 1 comment on HN

In plain words: Instead of guessing depth straight from pixels, ThirdEye feeds in separate frozen networks that each spot one visual cue—like occlusion edges, shading, or perspective—then blends them in stages, weighting each cue by how reliable it looks. Because those cue detectors stay fixed, the system reuses their training and needs only modest fine-tuning; no accuracy numbers are reported yet.

Abstract · THIRDEYE: Cue-Aware Monocular Depth Estimation via Brain-Inspired Multi-Stage Fusion

Monocular depth estimation methods traditionally train deep models to infer depth directly from RGB pixels. This implicit learning often overlooks explicit monocular cues that the human visual system relies on, such as occlusion boundaries, shading, and perspective. Rather than expecting a network to discover these cues unaided, we present ThirdEye, a cue-aware pipeline that deliberately supplies each cue through specialised, pre-trained, and frozen networks. These cues are fused in a three-stage cortical hierarchy (V1->V2->V3) equipped with a key-value working-memory module that weights them by reliability. An adaptive-bins transformer head then produces a high-resolution disparity map. Because the cue experts are frozen, ThirdEye inherits large amounts of external supervision while requiring only modest fine-tuning. This extended version provides additional architectural detail, neuroscientific motivation, and an expanded experimental protocol; quantitative results will appear in a future revision.

Calin Teodor Ioan
arXiv:2506.20877 · cs.CV, cs.AI · submitted Jun 25, 2025
abstract · pdf · html

add comment on HN

Would appreciate any form of peer feedback. Very few good MDE models. Trying to emulate how the brain would make use of geometric non-binocular features(occlusion, texture, priors/multi-frame, etc.), and spoon-feed it to a model. Then attempt to generate depth. Will post update once I have results.