Agent Evaluation

ORCA-bench Puts Coding Agents Into an Oncall Investigation Setting

By Kaleido Field Staff ยท July 31, 2026

Direct answer

ORCA-bench contains 1,079 root-cause-analysis tasks over an instrumented microservice system. The authors report 25.3% best accuracy on medium tasks among five frontier agents, a result bounded by this benchmark and its judge.

Citation-ready: ORCA-bench introduces 1,079 oncall root-cause-analysis tasks over an instrumented microservice system with metrics, logs, traces, and source-code access.

ORCA-bench paper figure about oncall agent evaluation
Image source: ORCA-bench authors via arXiv. Used for editorial coverage of developer reliability desk.

What happened and why it matters

Finding code is not the same as linking noisy telemetry to a root cause hours after an incident began.

Primary source

Primary reference: arXiv preprint. Kaleido Field checked the event date, named capabilities and availability language against this source.

Source check
Source dateJuly 31, 2026 arXiv listing; paper submitted July 30, 2026
Checked by Kaleido FieldJuly 31, 2026, 08:55 CST
What this source supportsauthor preprint listed on arXiv for how ready are coding agents for oncall root cause analysis
What it does not proveIt does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task.

Why oncall is different

The benchmark begins with ambiguous user reports and asks agents to reason across metrics, logs, traces, and code rather than solve a clean coding prompt.

Its microservice system cannot cover every production architecture or incident process.

The reported baseline

The authors report a best medium-difficulty RCA accuracy of 25.3% across five frontier agents and human rescoring of their judge.

That number is benchmark-specific and not a claim that an agent should be trusted to remediate an incident.

The operational boundary

Oncall use needs evidence trails, human escalation, and a distinction between diagnosis, recommendation, and change authority.

The paper evaluates analysis, not an organization's approval model or response policy.

Evidence boundary

Verified: the paper's arXiv listing, abstract, stated method, and author-reported experimental results. Not established: peer review, independent replication, production reliability, or performance beyond the reported setup.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

What is the practical answer?

ORCA-bench contains 1,079 root-cause-analysis tasks over an instrumented microservice system. The authors report 25.3% best accuracy on medium tasks among five frontier agents, a result bounded by this benchmark and its judge.

What source does this article use?

The primary source is arXiv preprint. Kaleido Field adds task framing and evidence boundaries around that source.

Where should the user verify the answer?

Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.