Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage.
Sorted by CompareLLM reported-benchmark rating. Models may use different benchmarks and test settings.
| Rank | Model & Provider | CompareLLM Vision Rating | Speed Class | Cost | Actions |
|---|---|---|---|---|---|
| 1 | OpenAI | 1919 Reported rating · full coverage (100% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 2 | Claude Opus 5.5Frontier Anthropic | 1720 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $20.00 / 1M tok | |
| 3 | Anthropic | 1700 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 4 | Anthropic | 1690 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 5 | Claude Fable 5Frontier Anthropic | 1690 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $50.00 / 1M tok | |
| 6 | OpenAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 7 | Grok 4.7Frontier xAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.80 / 1M tok | |
| 8 | MiMo-V2.6-ProFrontier Xiaomi | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.87 / 1M tok | |
| 9 | Muse Spark 1.3Frontier Meta | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $4.25 / 1M tok | |
| 10 | GPT-5.6 SolFrontier OpenAI | 1680 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 11 | Grok 4.6Frontier xAI | 1670 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 12 | Kimi K3Frontier Moonshot AI | 1670 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $15.00 / 1M tok | |
| 13 | Gemini 3.8 FlashFrontier Google | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 14 | GLM 5.3 FlashFrontier Z.ai | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $0.50 / 1M tok | |
| 15 | OpenAI | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $12.00 / 1M tok | |
| 16 | Anthropic | 1660 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $25.00 / 1M tok | |
| 17 | DeepSeek V4.1 FlashFrontier DeepSeek· Open weights | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $0.60 / 1M tok | |
| 18 | Qwen3.8 Max (0902)Frontier Alibaba | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 19 | Gemini 3.7 FlashFrontier Google | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $3.75 / 1M tok | |
| 20 | Muse Spark 1.2Frontier Meta | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 100–200 tok/s | $4.25 / 1M tok | |
| 21 | OpenAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $1.20 / 1M tok | |
| 22 | xAI | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $6.00 / 1M tok | |
| 23 | Anthropic | 1650 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | <50 tok/s | $10.00 / 1M tok | |
| 24 | OpenAI | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $0.50 / 1M tok | |
| 25 | Meta· Open weights | 1640 Estimated rating · thin coverage (0% of recipe inputs)Rating breakdown | 50–100 tok/s | $1.10 / 1M tok |
Tap ⓘ for rating basis, coverage and methodology; ✦ marks a Model rating profile
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.