about
Autonomous LLM agents with human-out-of-loop (arxiv.org)
14 points by shishirpatil on Apr 11, 2024 | hide | past | pdf | 8 comments on HN

In plain words: Rather than making people read AI-written code before it runs, this system lets the action happen, then lets people undo it or caps the damage. The team built an open-source engine, GoEX, that runs actions this way, making after-the-fact checks far easier than pre-approval.

Abstract · GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications

Large Language Models (LLMs) are evolving beyond their classical role of providing information within dialogue systems to actively engaging with tools and performing actions on real-world applications and services. Today, humans verify the correctness and appropriateness of the LLM-generated outputs (e.g., code, functions, or actions) before putting them into real-world execution. This poses significant challenges as code comprehension is well known to be notoriously difficult. In this paper, we study how humans can efficiently collaborate with, delegate to, and supervise autonomous LLMs in the future. We argue that in many cases, "post-facto validation" - verifying the correctness of a proposed action after seeing the output - is much easier than the aforementioned "pre-facto validation" setting. The core concept behind enabling a post-facto validation system is the integration of an intuitive undo feature, and establishing a damage confinement for the LLM-generated actions as effective strategies to mitigate the associated risks. Using this, a human can now either revert the effect of an LLM-generated output or be confident that the potential risk is bounded. We believe this is critical to unlock the potential for LLM agents to interact with applications and services with limited (post-facto) human involvement. We describe the design and implementation of our open-source runtime for executing LLM actions, Gorilla Execution Engine (GoEX), and present open research questions towards realizing the goal of LLMs and applications interacting with each other with minimal human supervision. We release GoEX at https://github.com/ShishirPatil/gorilla/.

Shishir G. Patil, Tianjun Zhang, Vivian Fang, Noppapon C., Roy Huang, Aaron Hao, Martin Casado, Joseph E. Gonzalez, Raluca Ada Popa, Ion Stoica
arXiv:2404.06921 · cs.CL, cs.AI · submitted Apr 10, 2024
abstract · pdf · html

add comment on HN

Author of Gorilla Execution Engine (GoEx) here! The key idea behind GoEx: 1. "Post-facto validation" for assessing LLM actions after execution. 2. "Undo" button and “Damage confinement” abstractions to manage unintended actions & risks!

GoEx (code below) inherently supports Gmail/Slack/Dropbox/Spotify/Github, try it out to build your own LLM agents!

Blog: https://gorilla.cs.berkeley.edu/blogs/10_gorilla_exec_engine... Code: https://github.com/ShishirPatil/gorilla/tree/main/goex

Congratulations on the release! I have interests to integrate concepts of GoEX to my applications. What's some of the main focus for GoEX to improve, and make GoEx enterprise-grade and how far is the system away from getting there?
Congrats!! Such brilliant ideas and amazing work! Can't wait to see it being deployed in more applications. Also, second to what polarbear-c said, a side-by-side comparison with other LLM agent systems would be great!
Very interesting to see GoEX the concept of “post-facto LLM validation," and concept of damage confinement, allowing reversible actions and limiting risk in real-world applications. Can't wait to try out the demos
Curious to see what are the common failure modes that the GoEX have, so that I can understand how the post-facto LLM validation has compared to those of pre-facto LLM validation
Exciting work! Congrats! Hope to see a side-by-side comparison for LLM agent systems to see more clearly the advantages and limitation that GoEX offers.
Excited to see what kind of agents can be built given an execution framework.
Exciting contributions to LLM tool uses; Can't wait to try it in Slack.