Optimized rankings calculated from CompareLLM category ratings Pro, LiveBench, and real-time token pricing economics.
Models ranked for repo work by our Coding rating, then price and speed. Our verdict, not a single lab score.
Lowest list-price models that still clear a usability bar. For high-volume chat and batch jobs.
Models ranked by dated output tokens-per-second and token generation speed. For long completions, batch generation, and agent loops that stream a lot of text.
Open-weight models ranked by our Reasoning rating. Check the license before you ship.
Models ranked for large document comprehension, massive codebase analysis, and RAG retrieval using verified context windows and input pricing.
Top reasoning models ranked by our Reasoning rating across all frontier architectures.
Context window and input list price first. For stuffing corpora, not for coding agents.
Our Agents rating for tool-using and computer-use stacks.
Models marked multimodal, ranked on Elo and latency for screenshot and document jobs.
Preference Elo first for long-form writing. Price still matters if you generate all day.
Workhorse models under $12/1M output. For daily coding and chat without Opus or Sol invoices.
Models that still clear a high preference-Elo bar, ranked so list price hurts.
Open-weight rows you can self-host. Not a laptop VRAM guide — check the license and your box.
Low list-price hosted models if Opus/Sonnet is too expensive. Then open the vs page against Sonnet 5.
Repo agents ranked by our Coding rating under a mid-tier budget. For CI bots that cannot burn Opus prices.
General assistants ranked for conversation quality. Elo first, then price if you chat all day.
Current top Anthropic vs top OpenAI model: our category ratings, speed and price.
Current top Google vs top OpenAI model on our category ratings and list price.
Current top xAI vs top OpenAI row. Price and Elo, not a personality contest.
Current top Anthropic vs top Google model. Coding, Elo, and output price.
Current top DeepSeek vs top Anthropic. The usual cheap-vs-frontier question.
Current top Meta open-weight row versus current top Anthropic row: our category ratings and price.
No single benchmark tells the whole story. Quality lists use CompareLLM category ratings with disclosed coverage. Price and speed stay exact facts and classes, never a blended value score.
Matchup hubs (like Claude vs GPT) evaluate the current top model from each provider on CompareLLM category ratings. As new flagships launch, rankings remap automatically without manual edits.
Rankings are computed from public benchmark snapshots and live pricing. Where we add an editorial recommendation it is labelled, dated, and attributed to CompareLLM — and it never reorders the computed table. No provider can pay for placement, for an editorial pick, or to boost their rank. We take no money from model providers.
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.