Enterprise Agents
WorkSurface-Bench Separates Enterprise Agent Routing From Retrieval
WorkSurface-Bench evaluates what its authors call surface routing across documents, tables, graphs, and cross-surface questions. Its auditable answer design is notable, but it remains a new benchmark rather than proof of enterprise-agent performance.
Citation-ready: WorkSurface-Bench introduces 1,151 atomic tasks for evaluating whether enterprise agents route questions to documents, tables, graphs, or cross-surface evidence.

What happened and why it matters
A correct answer can depend on selecting the right evidence surface before the agent retrieves or calls a tool.
Primary source
Primary reference: arXiv preprint. Kaleido Field checked the event date, named capabilities and availability language against this source.
| Source date | July 29, 2026 arXiv listing; paper submitted July 28, 2026 |
|---|---|
| Checked by Kaleido Field | July 30, 2026, 08:45 CST |
| What this source supports | author preprint listed on arXiv for what does WorkSurface-Bench measure for enterprise agents |
| What it does not prove | It does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task. |
The question before retrieval
The authors distinguish surface routing from retrieval because a narrative fact, a calculation, and a dependency relationship can require different evidence stores.
The benchmark's constructed workspaces do not represent every enterprise data model or access policy.
How answers are meant to be auditable
The paper says table answers are reproduced with executed DuckDB queries, document answers point to verified spans, and graph answers trace source dependencies.
Auditable references in a benchmark do not automatically carry over to a company's production connectors.
A useful deployment check
Ask whether an agent can name the evidence surface it selected and expose the query, passage, or graph path that supports its answer.
The study does not establish compliance, privacy, or permissions outcomes for a real workspace.
Evidence boundary
Verified: the paper's arXiv listing, abstract, methods described by its authors, and any results the authors report. Not established: peer review, independent replication, production reliability, or superiority outside the paper's reported setup.
FAQ
What is the practical answer?
WorkSurface-Bench evaluates what its authors call surface routing across documents, tables, graphs, and cross-surface questions. Its auditable answer design is notable, but it remains a new benchmark rather than proof of enterprise-agent performance.
What source does this article use?
The primary source is arXiv preprint. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.