about
ICRA: We Turned iPhones into Spatial Data Capture Rigs (arxiv.org)
2 points by satpal 138 days ago | hide | past | pdf | discuss on HN

In plain words: Ordinary smartphones can record hour-long first-person videos with continuous pose tracking, unlike today's few-minute clips, and a free pipeline turns the captures into training-ready data. Training a robot-control model on the resulting 200 hours lowered its action-prediction error on held-out data.

Abstract · MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes only a few minutes long, which fails to capture the long-horizon temporal dependencies that complex robotic task execution requires. We present MobileEgo Anywhere, a framework for collecting hour-plus egocentric trajectories on commodity mobile hardware that uses modern smartphone sensors for long-term pose tracking without the hardware barriers of traditional robotics data collection. We release three components: (1) STERA, an open-source video-processing pipeline that converts raw mobile captures into standardized, training-ready formats for VLA and foundation-model research; (2) a free mobile app that lets any user record egocentric activity; and (3) a 200-hour dataset of diverse, long-form egocentric data with persistent state tracking across 584 sessions. We further show this data is a usable training signal:mid-training a VLA on it lowers held-out action-prediction error.

Senthil Palanisamy, Abhishek Anand, Satpal Singh Rathore, Pratyush Patnaik, Shubhanshu Khatana, Ekaksh Janweja
arXiv:2605.05945 · cs.CV, cs.CL · submitted May 7, 2026 · updated Jul 8, 2026
abstract · pdf · html

add comment on HN