Models ranked for repo work by our Coding rating, then price and speed. Our verdict, not a single lab score.
Teal ring = metrics this list is ranked on.
| Rank | Model Name | Suitability | Coding Elo | Out $/1M | tok/s | Snappier | Compare |
|---|---|---|---|---|---|---|---|
| #1 | 66.0 | 1660 | $20/1M tok | 50–100 tok/s | 2s+ | 👑 Leader | |
| #2 | 66.0 | 1660 | $50/1M tok | <50 tok/s | 2s+ | vs Claude Opus 5.5 | |
| #3 | 64.1 | 1641 | $6/1M tok | <50 tok/s | 500ms–2s | vs Claude Opus 5.5 |
If Claude Opus 5.5 misses a constraint on Coding Elo, GPT-6 Astra is the next suitability-ranked option.
Ranking-grade snapshots only. Missing metrics stay neutral and models need at least 40% weighted evidence.
U = published provisionally after admin review; final CompareLLM verification is pending.
Software engineering tasks require deep repo-level reasoning, multi-file AST understanding, and test-driven patch synthesis rather than simple one-liner generation.
Avoiding hallucinations in complex imports, adhering strictly to existing codebase architectures, and resolving subtle regression bugs.
Ranked on CompareLLM Coding rating from disclosed recipe inputs (SWE-bench Pro, LiveCodeBench, SciCode). Missing inputs stay missing.
Our disclosed category rating for software engineering
Named reference the Coding recipe may consider
Exact commercial fact shown beside the rating, not inside it
Plain-English methodology and leaderboard answers