about
Building Browser Agents: Architecture, Security, and Practical Solutions (arxiv.org)
2 points by aramvr 301 days ago | hide | past | pdf | 1 comment on HN

In plain words: A report on running a web-browsing agent finds that how it is built matters more than model smarts, and hidden trick text makes autonomous browsing unsafe. Narrow tools with code-enforced limits solved about 85% of 53 web challenges, versus roughly 50% for earlier agents.

Abstract

Browser agents enable autonomous web interaction but face critical reliability and security challenges in production. This paper presents findings from building and operating a production browser agent. The analysis examines where current approaches fail and what prevents safe autonomous operation. The fundamental insight: model capability does not limit agent performance; architectural decisions determine success or failure. Security analysis of real-world incidents reveals prompt injection attacks make general-purpose autonomous operation fundamentally unsafe. The paper argues against developing general browsing intelligence in favor of specialized tools with programmatic constraints, where safety boundaries are enforced through code instead of large language model (LLM) reasoning. Through hybrid context management combining accessibility tree snapshots with selective vision, comprehensive browser tooling matching human interaction capabilities, and intelligent prompt engineering, the agent achieved approximately 85% success rate on the WebGames benchmark across 53 diverse challenges (compared to approximately 50% reported for prior browser agents and 95.7% human baseline).

Aram Vardanyan
arXiv:2511.19477 · cs.SE · submitted Nov 22, 2025
abstract · pdf · html · 30 pages, 22 figures. Production architecture and benchmark evaluation of browser agents

add comment on HN

Arxiv paper proposing a practical architecture for building browser-based AI agents. Covers system design patterns, sandboxing and security considerations, and how to make agents robust to real-world web complexity. Includes evaluations on common browsing tasks and discusses deployment tradeoffs and open problems.