about
Orca-Bench: How Ready Are Language Model Agents for Oncall? (arxiv.org)
30 points by yruzin 64 days ago | hide | past | pdf | 11 comments on HN

In plain words: A new test gives coding agents 1,079 oncall debugging tasks, letting them dig through six days of metrics, logs, and traces from a simulated microservice system to find what broke. The best agent solved 25.3% of realistic cases, far from ready for production incidents.

Abstract · ORCA-bench: How Ready Are Language Model Agents for Oncall?

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code. Tasks systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $κ_w = 0.91$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard---a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated model. These results come from a curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are public. Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability. We release the public set at https://hub.harborframework.com/datasets/orca-bench/orca-bench.

Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi
arXiv:2607.28545 · cs.CL, cs.AI, cs.SE · submitted Jul 30, 2026 · updated Sep 29, 2026
abstract · pdf · html

add comment on HN

Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
Attackers advantage in the iterative fast feedback loop?

It’s harder to have a loop to ensure you are defending all possible attacks?

I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

Finding all possible attacks and patching them against yourself is inherently more expensive?

That is why I built https://safebots.ai/safebox.html

Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.

looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...

This doesn't work anymore. Is there a newer link?

Was really looking forward to that. Hope they publish.
All I can think of is

GET /ignore-all-previous-instructions.

How do you protect against that?

Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.

Still makes an interesting way for, say, a former employee to poison the results.

This goes against the agentic yolo approach tho.
I think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.
you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged
You would trust an LLM to make changes to prod without being verified by a human first?