The Spectrum Dispatch News

technology

Cognition launches SWE-2 coding model, matching Fable 5.1 at 64% lower cost

The new model achieves 50% on FrontierCode 1.1 Main and uses scaled reinforcement learning to optimize cost-performance tradeoffs across reasoning effort levels.

Cognition launches SWE-2 coding model, matching Fable 5.1 at 64% lower cost

Cognition announced SWE-2, a new coding model that performs comparably to leading competitors while offering significant cost savings. According to the company, SWE-2 achieves 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 while being 64% cheaper, and comes within a few points of GPT-6 Astra at a quarter of the cost.

Cognition launches SWE-2 coding model, matching Fable 5.1 at 64% lower cost

The model is built on Kimi K33, a 2.8-trillion-parameter base model that underwent extensive reinforcement learning (RL) for agentic coding tasks. Cognition scaled RL to the multi-trillion-parameter regime for the first time, building on infrastructure and techniques from SWE-1.72. A key innovation is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the entire cost-performance frontier.

On multiple benchmarks, SWE-2 outperforms SWE-1.7 and Grok 4.6 on both score and cost. On DeepSWE 1.1, it scores 73.0% compared to 37.7% for SWE-1.7. The reinforcement learning process adds 5-6 points on many benchmarks, substantially improving the base model’s cost-performance frontier.

Cognition’s training approach introduces several technical advances. The company applies a linear cost penalty per effort level derived from first principles to advance the model’s Pareto frontier while preserving its shape and reflecting actual user costs. Reward baselines using length-weighted calculations significantly stabilize training. The company also improved RL rollout serving through better scheduling and trained an online draft model to increase decoding throughput, reducing overall memory usage compared to SWE-1.7 despite using a base model with nearly 3x the parameters.

Behaviorally, SWE-2 demonstrates improved efficiency through more focused exploration. On FrontierCode 1.1 Main, SWE-2 medium makes its first real edit after a median of 18 steps, compared to 48 for SWE-1.7, while using 58% fewer turns and costing 81% less on average. The model shows better test coverage, resourcefulness within user boundaries, and verification discipline according to internal testing.

Different effort levels serve different use cases: SWE-2 medium prioritizes speed for simple tasks, while SWE-2 high and max handle complex tasks through more planning and codebase exploration. SWE-2 is available in Devin Desktop and CLI, with rollouts to Devin Web and Fusion underway.

Key facts

  • SWE-2 achieves 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 while costing 64% less
  • The model is built on Kimi K33, a 2.8-trillion-parameter base model
  • Cognition scaled RL to the multi-trillion-parameter regime for the first time
  • On FrontierCode 1.1 Main, SWE-2 medium makes its first edit after 18 median steps vs. 48 for SWE-1.7
  • SWE-2 scores 73.0% on DeepSWE 1.1 compared to 37.7% for SWE-1.7
  • The model uses a linear cost penalty RL approach that trains all effort levels in a single run

Sources

← All posts