Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage.
Sorted by CompareLLM reported-benchmark rating. Models may use different benchmarks and test settings.
| Rank | Model & Provider | CompareLLM Agents Rating | Speed Class | Cost | Actions |
|---|---|---|---|---|---|
| 1 | OpenAI | 1662 Reported rating · full coverage (100% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 2 | Claude Opus 5.5Frontier Anthropic | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $20.00 / 1M tok | |
| 3 | OpenAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 4 | Muse Spark 1.3Frontier Meta | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.25 / 1M tok | |
| 5 | Anthropic | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 6 | GPT-5.6 SolFrontier OpenAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 7 | Claude Fable 5Frontier Anthropic | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 8 | Grok 4.7Frontier xAI | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.80 / 1M tok | |
| 9 | MiMo-V2.6-ProFrontier Xiaomi | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.87 / 1M tok | |
| 10 | Gemini 3.8 FlashFrontier Google | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 11 | GLM 5.3 FlashFrontier Z.ai | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.50 / 1M tok | |
| 12 | GLM-5.3Frontier Z.ai· Open weights | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $4.40 / 1M tok | |
| 13 | Grok 4.6Frontier xAI | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 14 | Kimi K3Frontier Moonshot AI | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $15.00 / 1M tok | |
| 15 | OpenAI | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $12.00 / 1M tok | |
| 16 | Anthropic | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 17 | OpenAI | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $0.50 / 1M tok | |
| 18 | DeepSeek V4.1 FlashFrontier DeepSeek· Open weights | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $0.60 / 1M tok | |
| 19 | Qwen3.8 Max (0902)Frontier Alibaba | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 20 | Alibaba | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $2.50 / 1M tok | |
| 21 | Gemini 3.7 FlashFrontier Google | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 22 | Qwen3.8 2.4T A95BFrontier Alibaba | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 23 | Meta· Open weights | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $1.10 / 1M tok | |
| 24 | Muse Spark 1.2Frontier Meta | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $4.25 / 1M tok | |
| 25 | Google | 1630 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $3.75 / 1M tok |
Tap ⓘ for rating basis, coverage and methodology; ✦ marks a Model rating profile
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.