Cognition has released SWE-2, its latest coding model trained via reinforcement learning on top of Moonshot AI’s 2.8-trillion-parameter open model, Kimi K3. In company benchmarks, SWE-2 achieved a score of 50.0% on FrontierCode 1.1 Main, matching the performance of Anthropic’s Claude Fable 5.1 while cutting inference costs by 64%. The model is integrated directly into Cognition’s Devin platform across Desktop and CLI interfaces, with no plans for a standalone API or open-weights release.
SWE-2 introduces selectable reasoning-effort levels trained within a single RL run using Pareto-informed cost penalties. This training setup allows users to balance execution speed and cost against task complexity. On FrontierCode 1.1 Main, the medium-effort setting of SWE-2 outperformed Cognition’s previous SWE-1.7 model while requiring 58% fewer turns and reducing operational costs by 81%.
To optimize rollout performance and lower training-inference mismatches, Cognition implemented technical upgrades including DSpark speculative decoding, prefill request batching, and NVFP4/FP8 quantization-aware kernels. While SWE-2 leads on Terminal-Bench 2.1 and approaches GPT-6 Astra’s performance at a lower cost, Cognition noted a performance drop on Terminal-Bench 4, where it trails rival models by roughly 30 points.
Why it matters
Cognition’s SWE-2 achieves competitive coding agent accuracy at a 64% lower cost, highlighting the leverage of RL on top of open base models.
Multi-effort reinforcement learning allows developers to explicitly trade off inference cost, speed, and accuracy within a single model architecture.
Cognition keeps its proprietary model gated behind the Devin ecosystem, reinforcing product lock-in over raw API distribution.
Source: marktechpost.com



