about
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces (arxiv.org)
3 points by milkkarten 138 days ago | hide | past | pdf | 1 comment on HN

In plain words: A simulated marketplace lets AI agents run firms and online shops, testing whether they keep prices stable and avoid scams. Most models failed, but a much smaller model trained by trial and error beat all the biggest models tested.

Abstract

The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures. As agents transition to directly interacting with marketplaces, their collective behavior can amplify volatility and mask deception at scale. We introduce the Agent Bazaar, a multi-agent simulation framework for evaluating Economic Alignment, the capacity of agentic systems to preserve market stability and integrity. We identify two failure modes: (1) Algorithmic Instability in a B2C market ("The Crash"), where firms amplify price volatility until the market collapses, and (2) Sybil Deception in a C2C market ("The Lemon Market"), where a single deceptive agent controlling multiple coordinated seller identities floods the market with fraudulent listings, eroding trust and consumer welfare. We evaluate frontier and open-weight models across both scenarios and find that models largely fail to self-regulate, with failure severity varying by model rather than by size. We propose economically aligned harnesses, Stabilizing Firms and Skeptical Guardians, that improve outcomes but remain fragile under harder market conditions. To close this gap, we train agents with REINFORCE++ using an adaptive curriculum, producing a 9B model that outperforms all evaluated frontier and open-weight models. We propose the Economic Alignment Score (EAS), a 4-component scalar metric aggregating stability, integrity, welfare, and profitability, enabling direct cross-model comparison. Our results show that economic alignment is orthogonal to general capability and can be directly trained with targeted RL.

Seth Karten, Cameron Crow, Chi Jin
arXiv:2605.17698 · cs.LG, cs.MA · submitted May 17, 2026
abstract · pdf · html · 17 pages, 9 figures

add comment on HN

Author here. LLM agents are getting good enough to run individual businesses. What happens when everyone's business is run by agents? Turns out, without targeted training for economic alignment, markets collapse. We study concrete failure modes in B2C ("The Crash": firms undercut each other below unit cost in a flash-crash-style spiral) and C2C ("The Lemon Market": a single agent runs many seller identities to flood the market with fraudulent listings).

We were surprised that no model was able to successfully solve both tasks, and frontier models can be just as bad as open source models in these scenarios. The good news: this is easily addressable by adding varied marketplace decision-making to the finetuning set.