Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage.
Sorted by CompareLLM reported-benchmark rating. Models may use different benchmarks and test settings.
| Rank | Model & Provider | CompareLLM Coding Rating | Speed Class | Cost | Actions |
|---|---|---|---|---|---|
| 1 | Anthropic | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $20.00 / 1M tok | |
| 2 | GPT-6 AstraFrontier OpenAI | 1660 Reported rating · full coverage (100% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 3 | Grok 4.6Frontier xAI | 1641 Reported rating · thin coverage (65% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 4 | Anthropic | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 5 | Claude Fable 5Frontier Anthropic | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 6 | OpenAI | 1620 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 7 | Muse Spark 1.3Frontier Meta | 1620 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.25 / 1M tok | |
| 8 | Anthropic | 1620 Reported rating · partial coverage (75% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 9 | GPT-5.6 SolFrontier OpenAI | 1620 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 10 | Grok 4.7Frontier xAI | 1610 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.80 / 1M tok | |
| 11 | MiMo-V2.6-ProFrontier Xiaomi | 1610 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.87 / 1M tok | |
| 12 | Kimi K3Frontier Moonshot AI | 1610 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $15.00 / 1M tok | |
| 13 | GLM 5.3 FlashFrontier Z.ai | 1600 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.50 / 1M tok | |
| 14 | OpenAI | 1600 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $12.00 / 1M tok | |
| 15 | Anthropic | 1600 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 16 | DeepSeek V4.1 FlashFrontier DeepSeek· Open weights | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $0.60 / 1M tok | |
| 17 | Qwen3.8 Max (0902)Frontier Alibaba | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 18 | Gemini 3.7 FlashFrontier Google | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 19 | Qwen3.8 2.4T A95BFrontier Alibaba | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 20 | Muse Spark 1.2Frontier Meta | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $4.25 / 1M tok | |
| 21 | xAI | 1590 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 22 | OpenAI | 1580 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $0.50 / 1M tok | |
| 23 | OpenAI | 1580 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $1.20 / 1M tok | |
| 24 | Anthropic | 1580 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 25 | Alibaba | 1570 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $2.50 / 1M tok |
Tap ⓘ for rating basis, coverage and methodology; ✦ marks a Model rating profile
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.