Cognition has announced the launch of its new SWE-2 model, marking a significant milestone in the development of artificial intelligence. The SWE-2 model is the closest yet to the frontier, outperforming its predecessor, SWE-1.7, and matching the capabilities of top models like GPT-5.6 Sol and Fable 5/5.1 at a fraction of their cost.
The SWE-2 model's advancements are attributed to its post-training from Kimi K33, a 2.8T-parameter model that underwent extensive reinforcement learning (RL) for agentic coding. This training enabled the model to achieve substantial headroom, adding 5-6 points on many benchmarks and shifting the cost-performance frontier.
One of the key features of the SWE-2 model is its ability to apply a linear cost penalty per effort level in a single RL run. This approach, derived from first principles, advances the model's entire Pareto frontier while preserving its shape and reflecting actual user costs in training. The model also utilizes a length-weighted reward baseline, which reduces gradient variance and stabilizes training.
In terms of behavioral patterns, the SWE-2 model exhibits improvements in test coverage, resourcefulness, and verification discipline. It is better at writing tests that check an implementation end-to-end, catching regressions and edge cases more reliably. The model is also more willing to look for alternative routes to the same answer when the obvious path is blocked.
The SWE-2 model's training data has been significantly improved, with the number of RL environments tripled and instruction-following overlays added. The model has also been trained to keep multiple instructions in context without losing sight of the underlying task. Additionally, the verifiers in the data have been hardened to prevent reward hacking.
In evaluations of model alignment and trustworthiness, the SWE-2 model passed 98.0% of attempts overall, with 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. The model also showed no statistically significant increase or decrease in vulnerability for any framing condition in the context-dependent vulnerability evaluation.
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
