Coding Models
Cognition Makes Reasoning Effort Central to SWE-2
Cognition introduced SWE-2 on September 10 and made it available in Devin Desktop and CLI, with other Devin surfaces rolling out. The company emphasizes training multiple reasoning-effort levels together. That makes the effort setting part of the model comparison, not a detail to omit beside a score or cost claim.
Citation-ready: Cognition says SWE-2 trains multiple reasoning-effort levels in one reinforcement-learning run and is available in Devin Desktop and CLI from September 10.
Evidence boundary: Developer release and self-reported experiments. No independent benchmark rerun or local inference trial; company charts are not imported as official third-party rankings.

What happened and why it matters
The release centers on the tradeoff between successful work and effort. An evaluation that mixes effort levels or ignores rejected outputs cannot establish that one model is cheaper for a team.
Primary evidence
Primary reference: Cognition SWE-2 technical release. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | September 10, 2026 |
|---|---|
| Checked by Kaleido Field | September 11, 2026, CST |
| Source function | coding agents -> cost per accepted task and evaluation provenance |
The training objective includes cost
Cognition describes SWE-2 as post-trained from Kimi K3 and applies an effort-specific cost penalty during reinforcement learning. The release also describes changes to serving and training environments. Those are the developer's account of the model, not an independently reproduced training result.
The important operational question is narrower: at a chosen effort level, can the agent complete the task and pass the relevant tests within an acceptable budget? More generated tokens or fewer tool calls alone do not answer that.
Keep unsuccessful runs in the denominator
For a team trial, record all attempts, including abandoned edits, retries and human corrections. Measure the accepted patch and its verification, then associate that outcome with the actual model, effort setting and tool environment.
A lower invoice for an unfinished patch is not lower cost per completed task. Conversely, a longer run can be worthwhile when it resolves a difficult failure that a quick attempt misses. This is an evaluation design, not a reported result for SWE-2.
Product access and benchmark standing are separate
Cognition names Desktop and CLI as available and describes Web and Fusion as rolling out. Check the selected surface rather than assuming that an announcement means simultaneous access everywhere.
Our open-model release evidence report makes a related distinction: a repository or release page establishes an artifact, while performance needs a specific test. We have not ranked SWE-2 against consumer visual apps.
Evidence boundary
Developer release and self-reported experiments. No independent benchmark rerun or local inference trial; company charts are not imported as official third-party rankings.
FAQ
Are Cognition's published charts independent evaluations?
Not on the evidence used here. They are results reported in the developer's own release and must retain that attribution.