about
Mosaic: A Modular System for Assistive and Interactive Cooking (arxiv.org)
2 points by rntn on Mar 16, 2024 | hide | past | pdf | discuss on HN

In plain words: A system that pairs big pre-trained language and vision models for understanding with small, precise controllers for robot motion, letting several robots cook alongside people and answer questions. In 60 full trials it handled complex cooking tasks involving multiple robots and humans.

Abstract · MOSAIC: Modular Foundation Models for Assistive and Interactive Cooking

We present MOSAIC, a modular architecture for coordinating multiple robots to (a) interact with users using natural language and (b) manipulate an open vocabulary of everyday objects. MOSAIC employs modularity at several levels: it leverages multiple large-scale pre-trained models for high-level tasks like language and image recognition, while using streamlined modules designed for low-level task-specific control. This decomposition allows us to reap the complementary benefits of foundation models as well as precise, more specialized models. Pieced together, our system is able to scale to complex tasks that involve coordinating multiple robots and humans. First, we unit-test individual modules with 180 episodes of visuomotor picking, 60 episodes of human motion forecasting, and 46 online user evaluations of the task planner. We then extensively evaluate MOSAIC with 60 end-to-end trials. We discuss crucial design decisions, limitations of the current system, and open challenges in this domain. The project's website is at https://portal-cornell.github.io/MOSAIC/

Huaxiaoyue Wang, Kushal Kedia, Juntao Ren, Rahma Abdullah, Atiksh Bhardwaj, Angela Chao, Kelly Y Chen, Nathaniel Chin, Prithwish Dan, Xinyi Fan, Gonzalo Gonzalez-Pumariega, Aditya Kompella, et al.
arXiv:2402.18796 · cs.RO · submitted Feb 29, 2024 · updated Oct 25, 2025
abstract · pdf · html · 22 pages, 13 figures; CoRL 2024

add comment on HN