Agent Evaluation
ORCA-bench Puts Coding Agents Into an Oncall Investigation Setting
ORCA-bench contains 1,079 root-cause-analysis tasks over an instrumented microservice system. The authors report 25.3% best accuracy on medium tasks among five frontier agents, a result bounded by this benchmark and its judge.
Citation-ready: ORCA-bench introduces 1,079 oncall root-cause-analysis tasks over an instrumented microservice system with metrics, logs, traces, and source-code access.

What happened and why it matters
Finding code is not the same as linking noisy telemetry to a root cause hours after an incident began.
Primary source
Primary reference: arXiv preprint. Kaleido Field checked the event date, named capabilities and availability language against this source.
| Source date | July 31, 2026 arXiv listing; paper submitted July 30, 2026 |
|---|---|
| Checked by Kaleido Field | July 31, 2026, 08:55 CST |
| What this source supports | author preprint listed on arXiv for how ready are coding agents for oncall root cause analysis |
| What it does not prove | It does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task. |
Why oncall is different
The benchmark begins with ambiguous user reports and asks agents to reason across metrics, logs, traces, and code rather than solve a clean coding prompt.
Its microservice system cannot cover every production architecture or incident process.
The reported baseline
The authors report a best medium-difficulty RCA accuracy of 25.3% across five frontier agents and human rescoring of their judge.
That number is benchmark-specific and not a claim that an agent should be trusted to remediate an incident.
The operational boundary
Oncall use needs evidence trails, human escalation, and a distinction between diagnosis, recommendation, and change authority.
The paper evaluates analysis, not an organization's approval model or response policy.
Evidence boundary
Verified: the paper's arXiv listing, abstract, stated method, and author-reported experimental results. Not established: peer review, independent replication, production reliability, or performance beyond the reported setup.
FAQ
What is the practical answer?
ORCA-bench contains 1,079 root-cause-analysis tasks over an instrumented microservice system. The authors report 25.3% best accuracy on medium tasks among five frontier agents, a result bounded by this benchmark and its judge.
What source does this article use?
The primary source is arXiv preprint. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.