Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage.
Sorted by CompareLLM reported-benchmark rating. Models may use different benchmarks and test settings.
| Rank | Model & Provider | CompareLLM Chat Rating | Speed Class | Cost | Actions |
|---|---|---|---|---|---|
| 1 | Anthropic | 1720 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $20.00 / 1M tok | |
| 2 | GPT-6 AstraFrontier OpenAI | 1700 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 3 | Anthropic | 1700 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 4 | Anthropic | 1690 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 5 | Claude Fable 5Frontier Anthropic | 1690 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 6 | OpenAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 7 | Grok 4.7Frontier xAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.80 / 1M tok | |
| 8 | MiMo-V2.6-ProFrontier Xiaomi | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.87 / 1M tok | |
| 9 | Muse Spark 1.3Frontier Meta | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.25 / 1M tok | |
| 10 | GPT-5.6 SolFrontier OpenAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 11 | GLM-5.3Frontier Z.ai· Open weights | 1670 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $4.40 / 1M tok | |
| 12 | Grok 4.6Frontier xAI | 1670 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 13 | Kimi K3Frontier Moonshot AI | 1670 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $15.00 / 1M tok | |
| 14 | Gemini 3.8 FlashFrontier Google | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 15 | GLM 5.3 FlashFrontier Z.ai | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.50 / 1M tok | |
| 16 | OpenAI | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $12.00 / 1M tok | |
| 17 | Anthropic | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 18 | DeepSeek V4.1 FlashFrontier DeepSeek· Open weights | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $0.60 / 1M tok | |
| 19 | Qwen3.8 Max (0902)Frontier Alibaba | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 20 | Gemini 3.7 FlashFrontier Google | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 21 | Qwen3.8 2.4T A95BFrontier Alibaba | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 22 | Muse Spark 1.2Frontier Meta | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $4.25 / 1M tok | |
| 23 | OpenAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $1.20 / 1M tok | |
| 24 | xAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 25 | Anthropic | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok |
Tap ⓘ for rating basis, coverage and methodology; ✦ marks a Model rating profile
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.