Cursor’s proprietary Composer 2.5 model has been released, and its benchmarks are nearly indistinguishable from Anthropic’s Claude Opus 4.7. On Terminal-Bench 2.0, Composer 2.5 scored 69.3% versus Opus 4.7’s 69.4%. In SWE-Bench Multilingual, the gap was similarly narrow: 79.8% for Composer against 80.5% for Opus. On CursorBench v3.1, Composer 2.5 also performed competitively. The company has essentially built a model that goes head-to-head with Anthropic’s offering.
3mo
Cursor’s proprietary Composer 2.5 model has been released, and its benchmarks are nearly indistinguishable from Anthropic’s Claude Opus 4.7. On Terminal-Bench 2.0, Composer 2.5 scored 69.3% versus Opus 4.7’s 69.4%. In SWE-Bench Multilingual, the gap was similarly narrow: 79.8% for Composer against 80.5% for Opus. On CursorBench v3.1, Composer 2.5 also performed competitively. The company has essentially built a model that goes head-to-head with Anthropic’s offering.
3mo
Ancora nessun commento. Sii il primo!
Commenti
Ancora nessun commento. Sii il primo!