Compare AI Models With Real Benchmarks
Dated ratings and API-pricing snapshots for 65 tracked models from 71 published catalog records. We use sourced benchmark reports and provider pricing; we do not run SWE-bench ourselves. Compare coding, reasoning, speed, and price.
What do you need a model for?
Tap to switchWhy this pick and trade-offs
- ✓CompareLLM Coding rating 1660 (our opinion at reported configurations)
- ✓Massive 1M tokens context window for deep document ingestion
Trade-off: Proprietary closed weights: requires commercial cloud API access
How this is scored rating
Scored on this task's own published evidence — quality 60%, cost 20%, performance 20%. Weights are renormalised over the dimensions that have published values.
Quality carries the largest share of this score.
How this is scored rating
Scored on this task's own published evidence — quality 60%, cost 20%, performance 20%. Weights are renormalised over the dimensions that have published values.
Quality carries the largest share of this score.
How this is scored rating
Scored on this task's own published evidence — quality 60%, cost 20%, performance 20%. Weights are renormalised over the dimensions that have published values.
Quality carries the largest share of this score.
AI Model Leaderboards & Benchmark Rankings
Rankings across 65 models from 71 public catalog SKUs with live benchmarks and pricing.Rankings across 65 tracked models (50 LLMs · 7 Image · 4 Video) with live benchmarks and pricing.
How models are selected for each category13 Categories · Selection Rules
CompareLLM Coding rating at one effort. Default inputs are SWE-bench Pro, LiveCodeBench and SciCode. Missing inputs are skipped, never filled. Price and context never enter this list.
CompareLLM Reasoning rating at one effort. Default inputs are HLE and IFEval. GPQA and LiveBench are not dormant defaults.
CompareLLM Writing rating at one effort. Default input is IFEval. Facts stay beside the rating, not inside it.
CompareLLM Agents rating at one effort. Reference inputs can anchor; throughput and TTFT may adjust an already anchored rating. Facts alone cannot rate.
CompareLLM Chat rating at one effort. IFEval can anchor; TTFT, throughput and price may adjust. Votes matter more here because chat quality is preference.
For Gemini-class models that read images. Not image generation. Default input is MMMU-Pro.
CompareLLM Image generation Elo. T2I-CompBench++ can anchor; price and generation time may adjust. Compact speed remains a class.
CompareLLM Video generation Elo. Separate from Video understanding. VBench-2.0 can anchor.
No default reference in recipe-v1. Honest own/vote-only rating or unrated until a reference clears governance. Never borrows VBench generation quality.
CompareLLM Audio understanding Elo. Separate from ASR, TTS and music generation.
RTFx is a performance fact and cannot anchor Elo without own/votes or a later eligible reference. Not MMAU. Not inverse RTF.
No default reference in recipe-v1. Own/vote-only or unrated. Music is a different category.
No default reference in recipe-v1. Own/vote-only or unrated. Not ranked on TTS naturalness.
Compare AI Models Head-to-Head
Explore comparisons across 71 public models, with benchmark deltas, capability radar percentiles, and list-price cost projections.
Best Value Right Now
8 of 50 priced models sit on the current price-versus-Coding-rating frontier — nothing cheaper has a higher CompareLLM Coding rating, nothing stronger costs less.
What Changed Today
Latest verified model releases, benchmark shifts, and price moves.
AI Benchmark News & Analysis
Empirical breakdowns of newly released weights, price adjustments, and benchmark shifts.

Six New Models in Two Days: Opus 5.5, GPT-6 Sol and Luna, Grok 4.7, MiMo-V2.6-Pro, Command A+
Between 21 and 22 September, six models with a published Intelligence Index score reached OpenRouter. Claude Opus 5.5 takes the top of the index at 58. Xiaomi's MiMo-V2.6-Pro matches Grok 4.7 at a fifth of the output price, and Cohere's Command A+ returns a first token in 0.3 seconds.

GPT-6 Sol and Luna: Same Price and Faster for Sol, Half the Price for Luna
OpenAI filled out the GPT-6 line on 22 September. Sol keeps GPT-5.6 Sol's $2/$10 price, scores one point higher and streams 37% faster. Luna halves input cost and cuts output cost by 58%, for one point less on the Intelligence Index.

Grok 4.7: Two Points Up on the Index, 20% Cheaper, 27% Faster Than Grok 4.6
xAI's Grok 4.7 lists at $1.60 input and $4.80 output per million tokens, 20% below Grok 4.6. It streams 27% faster, and it edges up from 44 to 46 on the Artificial Analysis Intelligence Index. Context and output ceilings are unchanged at 500K and 450K.
Frequently Asked Questions
Methodology, metric definitions, and how CompareLLM scores models.
