Coding Models

Cognition Makes Reasoning Effort Central to SWE-2

By Kaleido Field Staff ยท September 11, 2026

Compare the setting as well as the model name

Cognition introduced SWE-2 on September 10 and made it available in Devin Desktop and CLI, with other Devin surfaces rolling out. The company emphasizes training multiple reasoning-effort levels together. That makes the effort setting part of the model comparison, not a detail to omit beside a score or cost claim.

Citation-ready: Cognition says SWE-2 trains multiple reasoning-effort levels in one reinforcement-learning run and is available in Devin Desktop and CLI from September 10.

Evidence boundary: Developer release and self-reported experiments. No independent benchmark rerun or local inference trial; company charts are not imported as official third-party rankings.

Cognition SWE-2 release artwork with a cost-performance plot
Image source: Cognition; developer release artwork, not an independent benchmark. Used for editorial coverage of coding agent evaluation desk.

What happened and why it matters

The release centers on the tradeoff between successful work and effort. An evaluation that mixes effort levels or ignores rejected outputs cannot establish that one model is cheaper for a team.

Primary evidence

Primary reference: Cognition SWE-2 technical release. Kaleido Field checked the event date and the article's attributed facts against this source.

Source check
Source dateSeptember 10, 2026
Checked by Kaleido FieldSeptember 11, 2026, CST
Source functioncoding agents -> cost per accepted task and evaluation provenance

The training objective includes cost

Cognition describes SWE-2 as post-trained from Kimi K3 and applies an effort-specific cost penalty during reinforcement learning. The release also describes changes to serving and training environments. Those are the developer's account of the model, not an independently reproduced training result.

The important operational question is narrower: at a chosen effort level, can the agent complete the task and pass the relevant tests within an acceptable budget? More generated tokens or fewer tool calls alone do not answer that.

Keep unsuccessful runs in the denominator

For a team trial, record all attempts, including abandoned edits, retries and human corrections. Measure the accepted patch and its verification, then associate that outcome with the actual model, effort setting and tool environment.

A lower invoice for an unfinished patch is not lower cost per completed task. Conversely, a longer run can be worthwhile when it resolves a difficult failure that a quick attempt misses. This is an evaluation design, not a reported result for SWE-2.

Product access and benchmark standing are separate

Cognition names Desktop and CLI as available and describes Web and Fusion as rolling out. Check the selected surface rather than assuming that an announcement means simultaneous access everywhere.

Our open-model release evidence report makes a related distinction: a repository or release page establishes an artifact, while performance needs a specific test. We have not ranked SWE-2 against consumer visual apps.

Evidence boundary

Developer release and self-reported experiments. No independent benchmark rerun or local inference trial; company charts are not imported as official third-party rankings.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

Are Cognition's published charts independent evaluations?

Not on the evidence used here. They are results reported in the developer's own release and must retain that attribution.