Agent Research
Desktop-Delta Bench Tests Whether Computer-Use Agents Track Screen Changes
Desktop-Delta Bench is a new author-released benchmark with 2,013 human-verified desktop-transition instances. It targets state verification and source tracking between actions; it is not evidence that any deployed agent is reliable on a user's computer.
Citation-ready: Desktop-Delta Bench, listed on arXiv on July 29, contains 2,013 human-verified examples for evaluating whether computer-use models identify task-relevant desktop GUI transitions.

What happened and why it matters
A GUI agent can finish a task only if it can distinguish a real state change from a delayed or irrelevant screenshot.
Primary source
Primary reference: arXiv preprint. Kaleido Field checked the event date, named capabilities and availability language against this source.
| Source date | July 29, 2026 arXiv listing; paper submitted July 28, 2026 |
|---|---|
| Checked by Kaleido Field | July 30, 2026, 08:45 CST |
| What this source supports | author preprint listed on arXiv for what does Desktop-Delta Bench measure for computer-use agents |
| What it does not prove | It does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task. |
The missing unit of evaluation
The authors argue that end-task success and single-frame grounding can miss whether an agent correctly reconstructed the transition caused by its own action.
That is a benchmark claim about a controlled evaluation task, not a reliability finding for a particular desktop agent.
Why asynchronous screens matter
Remote input, application rendering, and screenshot capture can be delayed or unrelated to the prior action. The benchmark isolates state verification, source tracking, and action attribution as failure dimensions.
A model can still fail on an unrepresented application, timing pattern, or permission model.
What an operator can take from it
For an agent that changes files, forms, or settings, capture an explicit before-and-after state and require a check tied to the intended outcome.
The paper does not prescribe production controls or establish that its benchmark predicts real-world incident rates.
Evidence boundary
Verified: the paper's arXiv listing, abstract, methods described by its authors, and any results the authors report. Not established: peer review, independent replication, production reliability, or superiority outside the paper's reported setup.
FAQ
What is the practical answer?
Desktop-Delta Bench is a new author-released benchmark with 2,013 human-verified desktop-transition instances. It targets state verification and source tracking between actions; it is not evidence that any deployed agent is reliable on a user's computer.
What source does this article use?
The primary source is arXiv preprint. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.