Top reasoning models ranked by our Reasoning rating across all frontier architectures.
Teal ring = metrics this list is ranked on.
| Rank | Model Name | Suitability | Reasoning Elo | Out $/1M | tok/s | Snappier | Compare |
|---|---|---|---|---|---|---|---|
| #1 | 81.4 | 1892 | $50/1M tok | <50 tok/s | 2s+ | 👑 Leader | |
| #2 | 69.0 | 1680 | $0.87/1M tok | <50 tok/s | 2s+ | vs GPT-6 Astra | |
| #3 | 67.8 | 1660 | $0.5/1M tok | <50 tok/s | 2s+ | vs GPT-6 Astra |
If GPT-6 Astra misses a constraint on Reasoning Elo, MiMo-V2.6-Pro is the next suitability-ranked option at $0.87/1M tok output.
Ranking-grade snapshots only. Missing metrics stay neutral and models need at least 40% weighted evidence.
U = published provisionally after admin review; final CompareLLM verification is pending.
Evaluating the optimal model for "best reasoning llm" requires balancing domain capability, inference economics, and prompt adherence.
Navigating trade-offs between peak frontier capability, latency constraints, and operational API costs.
Ranked on the CompareLLM category rating that matches the query, with exact commercial facts and speed classes shown beside it.
Disclosed recipe for the matching job, with coverage
Recipe inputs such as SWE-bench Pro stay labelled as references
Exact commercial facts, never mixed into the rating
Plain-English methodology and leaderboard answers