CompareLLM
CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Updated Aug 16, 2026·100 Models·Live Ingest

Compare AI Models on Verified Benchmarks, Speed & Real Cost

Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.

Best codingBest cheap
Intelligence vs API Price Tradeoff

Quality vs. Price Pareto Frontier

Top-Left = Best Value

Plotted by intelligence (Elo) vs cost (Out $). The teal line connects models with unbeatable price-to-performance.

Leaderboard TableTable↓
30/71 models shown
Better
↑
Crowd vote (Elo)↓
Lower
↑ Up: Crowd vote (Elo)← Left: Cheaper

Vertical ↑

Crowd vote (Elo)

Lost more votes ↓People liked it more ↑

Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.

Horizontal →

Output (answer) price / 1 million tokens (USD)

← CheaperMore expensive →

Cost to generate the answer. Provider list prices are USD

Best-Value Frontier Line (7 Models)Models defining the teal efficiency boundary
Brand:
💬 LLMs & ReasoningLLMs🎨 Text-to-Image ModelsText-to-ImageNew
Sortable LLM Leaderboard

Live AI Model Leaderboard

Top 10 of 13 models.Top 13 of 13 models. Sortable rankings from dated verified benchmarks.

Compare every row vs:
All ModelsFrontierOpen Weights🇨🇳 China
1GLM-5.3🇨🇳 CNFrontier

Zhipu

Elo1,558
SWE76.4%
In /1M$0.5
Out /1M$1.8
vs Claude Opus 5Specs
2DeepSeek V4 Pro🇨🇳 CNOpenFrontier

DeepSeek

Elo1,536
SWE71.6%
In /1M$0.4
Out /1M$1.6
vs Claude Opus 5Specs
3Kimi K3🇨🇳 CNFrontier

Moonshot

Elo1,528
SWE73.8%
In /1M$0.6
Out /1M$2.5
vs Claude Opus 5Specs
4Qwen 3 Max🇨🇳 CNFrontier

Alibaba

Elo1,518
SWE67.8%
In /1M$1.2
Out /1M$4.8
vs Claude Opus 5Specs
5MiniMax M2.5🇨🇳 CNOpenFrontier

MiniMax

Elo1,478
SWE75.8%
In /1M$0.3
Out /1M$1.2
vs Claude Opus 5Specs
6Doubao Pro 1.5🇨🇳 CNFrontier

ByteDance

Elo1,470
SWE62.1%
In /1M$0.11
Out /1M$0.28
vs Claude Opus 5Specs
7Ernie 4.5 Turbo🇨🇳 CN

Baidu

Elo1,450
SWE57%
In /1M$0.3
Out /1M$1.2
vs Claude Opus 5Specs
8Qwen 2.5 Plus🇨🇳 CN

Alibaba

Elo1,445
SWE59.2%
In /1M$0.25
Out /1M$1
vs Claude Opus 5Specs
9Yi-Large🇨🇳 CN

01.AI

Elo1,430
SWE53.8%
In /1M$0.3
Out /1M$0.3
vs Claude Opus 5Specs
10Qwen 2.5 Coder 32B🇨🇳 CNOpen

Alibaba

Elo1,425
SWE65.2%
In /1M$0.18
Out /1M$0.72
vs Claude Opus 5Specs
#Model Elo LiveBench SWE-bench tok/s In $/1M Out $/1M Compare vs
1
GLM-5.3🇨🇳 CNFrontier

Zhipu

1,55870.8%76.4%90 tok/s$0.50 / 1M$1.80 / 1M
2
DeepSeek V4 Pro🇨🇳 CNOpenFrontier

DeepSeek

1,53670.1%71.6%70 tok/s$0.40 / 1M$1.60 / 1M
3
Kimi K3🇨🇳 CNFrontier

Moonshot

1,52869.5%73.8%68 tok/s$0.60 / 1M$2.50 / 1M
4
Qwen 3 Max🇨🇳 CNFrontier

Alibaba

1,51869.4%67.8%80 tok/s$1.20 / 1M$4.80 / 1M
5
MiniMax M2.5🇨🇳 CNOpenFrontier

MiniMax

1,47864.4%75.8%90 tok/s$0.30 / 1M$1.20 / 1M
6
Doubao Pro 1.5🇨🇳 CNFrontier

ByteDance

1,47064%62.1%105 tok/s$0.11 / 1M$0.28 / 1M
7
Ernie 4.5 Turbo🇨🇳 CN

Baidu

1,45061%57%95 tok/s$0.30 / 1M$1.20 / 1M
8
Qwen 2.5 Plus🇨🇳 CN

Alibaba

1,44563.5%59.2%110 tok/s$0.25 / 1M$1.00 / 1M
9
Yi-Large🇨🇳 CN

01.AI

1,43060.5%53.8%90 tok/s$0.30 / 1M$0.30 / 1M
10
Qwen 2.5 Coder 32B🇨🇳 CNOpen

Alibaba

1,42564.8%65.2%110 tok/s$0.18 / 1M$0.72 / 1M
11
DeepSeek Coder V2🇨🇳 CNOpen

DeepSeek

1,36558.6%60.5%75 tok/s$0.14 / 1M$0.28 / 1M
12
DeepSeek R1🇨🇳 CNOpen

DeepSeek

1,35859%65.2%62 tok/s$0.55 / 1M$2.19 / 1M
13
DeepSeek V3🇨🇳 CNOpen

DeepSeek

1,31055.4%48.6%66 tok/s$0.27 / 1M$1.10 / 1M
13 models in this filter · sorted by elo
All benchmark leaderboardsIntent rankings
US vs China AI FrontierGap Closed & Surpassing on Cost

How Close is China's AI to US Flagships?

Quality is virtually tied on independent benchmarks, while open-weight Chinese frontiers offer substantial price-to-performance advantages.

🇺🇸 US LeaderClaude Opus 5
vs
🇨🇳 China RivalGLM-5.3
General IntelligenceChatbot Arena human preference score
🇺🇸 1624 Elo
vs
🇨🇳 1558 Elo
Near Parity (96%)
Software EngineeringSWE-bench real GitHub bug resolution
🇺🇸 79.2%
vs
🇨🇳 76.4%
Near Parity (96%)
API Pricing & CostOutput price per 1 Million tokens
🇺🇸 $25.00
vs
🇨🇳 $1.80
13.9× Cheaper
Model AvailabilityWeights download & self-hosting
🇺🇸 Closed API
vs
🇨🇳 Open Weights
Open Source
Key Takeaway:

GLM-5.3 matches 96% of Claude Opus 5's intelligence at 93% lower cost.

Full Showdown Fact Sheet

Leaders Matchup

Open Claude Opus 5 vs GPT-5.6 Sol fact sheetOpen showdown fact sheet|View all comparisons
Multi-Dimensional Capability Radar

Frontier Trio Capability Radar

Comparing top Western standard models with China's leading frontier rival across 6 skill dimensions. Tap any spoke or dot to inspect.

Tap any node to inspect
0–100 %ile
Verified Intelligence & Dispatches

Recent AI News & Benchmark Briefings

Independent analysis on model promotions, price reductions, and verified score movements.

View all dispatches (50)
Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama
compare
CompareLLM Intelligence DeskAug 16, 2026

Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama

$1/$5 hosted Claude vs a cheap long-context open sibling. The alternative-to-Sonnet pair that is honest.

Shipped & Updated

New and Refreshed Models

View all 100 catalog models
Zhipu2026-08-14

GLM-5.3

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

Alibaba2026-08-14

Qwen3.8 27B

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

How to read this leaderboard

Plain-English methodology and leaderboard answers

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.
Fastest
Opus 5 vs GPT-5.6
Open weights
Top Verified Frontiers
Top Preference Elo
1,624Claude Opus 5
Fastest TTFT
82 msGemini 3.7 Flash
Best Value (Elo ≥ 1450)
$0.14/1MYi-Lightning
Indexable Models
7163 Active

Head-to-Head Comparison Studio

All Comparisons Matrix
Trending Comparisons & RivalriesVerified Telemetry
Frontier Titan

Claude Opus 5 vs GPT-5.6 Sol

Top flagship duel for #1 general intelligence

East vs West

Claude Sonnet 5 vs DeepSeek V4

Western standard vs Chinese frontier price/perf leader

Open Weights

Qwen3.8 vs Llama 4 Scout

Top self-hostable open-weight models compared

High-Speed

GLM-5.3 vs Gemini 3.7 Flash

Fast reasoning & latency vs low $/1M API cost

Verified briefRead dispatch
Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE
launch
CompareLLM Intelligence DeskAug 14, 2026

Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE

Alibaba open-weight 27B drop. Local/open-weight lists should see it. It will not win frontier-agents.

Verified briefRead dispatch
GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base
launch
CompareLLM Intelligence DeskAug 14, 2026

GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base

Coding-plan live; open weights promised after a two-week safety review. Newest Zhipu row.

Verified briefRead dispatch
Google2026-08-13

Gemini 3.7 Flash

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

xAI2026-08-12

Grok 4.6

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

ByteDance2026-08-10

Seed 2.1 Turbo

ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.

DeepSeek2026-07-31

DeepSeek V4 Flash

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.