Skip to main content
AI Index· Audited Sep 23, 2026
65 tracked of 71

Compare AI Models With Real Benchmarks

What do you need a model for?

Tap to switch
Coding rating1660
Overall rating1700
Input $/1M
$4
Output $/1M
$20
Why this pick and trade-offs
  • ✓CompareLLM Coding rating 1660 (our opinion at reported configurations)
  • ✓Massive 1M tokens context window for deep document ingestion

Trade-off: Proprietary closed weights: requires commercial cloud API access

Active Leaderboard65 tracked of 71

AI Model Leaderboards & Benchmark Rankings

Rankings across 65 models from 71 public catalog SKUs with live benchmarks and pricing.

How models are selected for each category
Tracked sets include verified published models within administrative modality caps. Category pools filter exclusively by model capability and subtype. Ratings reflect declared evidence recipes with zero fabricated or imputed data.
Coding50/50 rated

CompareLLM Coding rating at one effort. Default inputs are SWE-bench Pro, LiveCodeBench and SciCode. Missing inputs are skipped, never filled. Price and context never enter this list.

SWE-bench ProLiveCodeBenchSciCode
Reasoning50/50 rated

CompareLLM Reasoning rating at one effort. Default inputs are HLE and IFEval. GPQA and LiveBench are not dormant defaults.

Humanity's Last ExamIFEval
Writing50/50 rated

CompareLLM Writing rating at one effort. Default input is IFEval. Facts stay beside the rating, not inside it.

IFEval

CompareLLM Agents rating at one effort. Reference inputs can anchor; throughput and TTFT may adjust an already anchored rating. Facts alone cannot rate.

SWE-bench ProLiveCodeBenchHLEOutput speedTime to first token
Chat & support50/50 rated

CompareLLM Chat rating at one effort. IFEval can anchor; TTFT, throughput and price may adjust. Votes matter more here because chat quality is preference.

IFEvalTime to first tokenOutput speedOutput price

For Gemini-class models that read images. Not image generation. Default input is MMMU-Pro.

MMMU-Pro

CompareLLM Image generation Elo. T2I-CompBench++ can anchor; price and generation time may adjust. Compact speed remains a class.

T2I-CompBench++Image priceGeneration time

CompareLLM Video generation Elo. Separate from Video understanding. VBench-2.0 can anchor.

VBench-2.0Video priceGeneration time

No default reference in recipe-v1. Honest own/vote-only rating or unrated until a reference clears governance. Never borrows VBench generation quality.

CompareLLM Audio understanding Elo. Separate from ASR, TTS and music generation.

MMAR

RTFx is a performance fact and cannot anchor Elo without own/votes or a later eligible reference. Not MMAU. Not inverse RTF.

ASR RTFx

No default reference in recipe-v1. Own/vote-only or unrated. Music is a different category.

No default reference in recipe-v1. Own/vote-only or unrated. Not ranked on TTS naturalness.

Category and Overall ratings represent CompareLLM opinions from declared evidence recipes. Commercial facts and speed classes remain separate.Methodology & Recipes →
Head-to-Head Matrix

Compare AI Models Head-to-Head

Explore comparisons across 71 public models, with benchmark deltas, capability radar percentiles, and list-price cost projections.

Best Value Right Now

8 of 50 priced models sit on the current price-versus-Coding-rating frontier — nothing cheaper has a higher CompareLLM Coding rating, nothing stronger costs less.

What Changed Today

Latest verified model releases, benchmark shifts, and price moves.

AI Benchmark News & Analysis

Empirical breakdowns of newly released weights, price adjustments, and benchmark shifts.

Frequently Asked Questions

Methodology, metric definitions, and how CompareLLM scores models.

A CompareLLM rating is our own verdict for one job — Coding, Reasoning, Writing and so on — at one reasoning effort. It runs roughly 1000 to 2000 and is computed from the inputs that category's published recipe names, so it moves when the evidence moves. It is not LMArena's preference Elo, which is a separate third-party number we may consider as one reference among several.