about
Magentic-UI: Towards Human-in-the-Loop Agentic Systems (arxiv.org)
38 points by fitzn on Jul 31, 2025 | hide | past | pdf | 9 comments on HN

In plain words: A free web interface keeps a person in charge while an AI agent browses, runs code, and edits files, with shared planning and safety checks on risky actions. Tests with users showed it handles complex tasks with little human effort and stops unsafe steps.

Abstract · Magentic-UI: Towards Human-in-the-loop Agentic Systems

AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-level performance in most domains including computer use, software development, and research. Their growing autonomy and ability to interact with the outside world, also introduces safety and security risks including potentially misaligned actions and adversarial manipulation. We argue that human-in-the-loop agentic systems offer a promising path forward, combining human oversight and control with AI efficiency to unlock productivity from imperfect systems. We introduce Magentic-UI, an open-source web interface for developing and studying human-agent interaction. Built on a flexible multi-agent architecture, Magentic-UI supports web browsing, code execution, and file manipulation, and can be extended with diverse tools via Model Context Protocol (MCP). Moreover, Magentic-UI presents six interaction mechanisms for enabling effective, low-cost human involvement: co-planning, co-tasking, multi-tasking, action guards, and long-term memory. We evaluate Magentic-UI across four dimensions: autonomous task completion on agentic benchmarks, simulated user testing of its interaction capabilities, qualitative studies with real users, and targeted safety assessments. Our findings highlight Magentic-UI's potential to advance safe and efficient human-agent collaboration.

Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, et al.
arXiv:2507.22358 · cs.AI, cs.HC · submitted Jul 30, 2025
abstract · pdf · html

add comment on HN

Looks clean and all, but is this “multitask on steroid” actually requested, or is it just another nice demo for how agents can perform tasks etc? I just feel stressed that I have multiple tasks ongoing that I would require to switch context between every time anything needs input or fails to deliver what I fundamentally requested.
This seems to be the wet dream of mgmt leader types. It's not enough that AI/LLM tools make tasks more efficient, or quicker - it's their people can then be just as busy with these tools.

My feeling is: buzz off, this is like orgs now where the quickest people just get more work piled on for little gain.

I gladly use tools that give me some breathing room, and gives me some time to think about improvements, the future, etc. But will actively drag my feet if mgmt starts mandating these tools be used to shovel more tasks, tickets, or busy work at me.

Agreed, and I guess we can only pray that I won’t come to that.

I want the tools to enhance my current workflow and make me better/more accurate at what I am good at. I feel that develops me as a person as well.

In other words a case for the "Intelligent Workspace":

https://news.ycombinator.com/item?id=44627910

In lieu of chatbots as the primary means of working with AI.

This is an approach that is human centered and intended to accommodate a wide array of possible use cases where human interaction/engagement is essential for getting work done.

Integrating human-in-loop tooling: https://youtu.be/srG5Ze7mS7s

There’s room in the world for a startup whose tech lets AI agents request tasks be done by a human.

Upon a request being sent, one human from an army of anonymous humans gets given say, access to the agent’s browser that’s stuck, plus a description of what to do (“finish booking this flight, here’s what my user asked of me”). And they do the job and click “completed” which hands back to the agent, then move on to their next assignment (or go idle in a queue)

Dystopian, but bound to happen.

AI offloading tasks to Mechanical Turk
Incredibly weak to not seize the opportunity to name this «Mangentic» :D