about
Benchmarking the continuous improvement of language agents in deployment (arxiv.org)
2 points by polymorph1sm on Jun 19, 2024 | hide | past | pdf | discuss on HN

In plain words: StreamBench feeds language agents a steady stream of tasks and feedback, then checks whether they get better over time instead of just testing what they already know. Tests of simple improvement strategies reveal which parts of learning from past feedback matter most.

Abstract · StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improve over time. To address this gap, we introduce StreamBench, a pioneering benchmark designed to evaluate the continuous improvement of LLM agents over an input-feedback sequence. StreamBench simulates an online learning environment where LLMs receive a continuous flow of feedback stream and iteratively enhance their performance. In addition, we propose several simple yet effective baselines for improving LLMs on StreamBench, and provide a comprehensive analysis to identify critical components that contribute to successful streaming strategies. Our work serves as a stepping stone towards developing effective online learning strategies for LLMs, paving the way for more adaptive AI systems in streaming scenarios. Source code: https://github.com/stream-bench/stream-bench. Benchmark website: https://stream-bench.github.io.

Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee
arXiv:2406.08747 · cs.CL · submitted Jun 13, 2024 · updated Oct 31, 2024
abstract · pdf · html · NeurIPS 2024 Track on Datasets and Benchmarks

add comment on HN