CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Updated Aug 16, 2026·100 Models·Live Ingest

Compare AI Models on Verified Benchmarks, Speed & Real Cost

Independent crowd preference Elo, SWE-bench coding tests, and live $/1M API pricing with zero synthetic bias.

Best codingBest cheap
Intelligence vs API Price Tradeoff

Quality vs. Price Pareto Frontier

Top-Left = Best Value

Plotted by intelligence (Elo) vs cost (Out $). The teal line connects models with unbeatable price-to-performance.

Leaderboard TableTable↓
30/71 models shown
Better
↑
Crowd vote (Elo)↓
Lower
↑ Up: Crowd vote (Elo)← Left: Cheaper

Vertical ↑

Crowd vote (Elo)

Lost more votes ↓People liked it more ↑

Two hidden answers. A person picks the one they like. Higher Elo means more wins — not a school test or a coding exam.

Horizontal →

Output (answer) price / 1 million tokens (USD)

← CheaperMore expensive →

Cost to generate the answer. Provider list prices are USD

Best-Value Frontier Line (7 Models)Models defining the teal efficiency boundary
Brand:
💬 LLMs & ReasoningLLMs🎨 Text-to-Image ModelsText-to-ImageNew
Text-to-Image Benchmark Ranking

Top Text-to-Image Models (Arena Elo & Speed)

Top 4 of 4 models.Top 4 of 4 models. Rankings from blind preference votes, generation latency, and price per 1k images.

Compare every row vs:
All ModelsFrontierOpen Weights🇨🇳 China
1Wan2.6 Text to Image🇨🇳 CNOpen

Alibaba

Img Elo1,215
Latency3.1s
$/1k$30/1k
Prompt %89.5%
vs GPT Image 2 (High)Specs
2Seedream 4.0🇨🇳 CN

ByteDance

Img Elo1,225
Latency2.8s
$/1k$30/1k
Prompt %89%
vs GPT Image 2 (High)Specs
3Qwen Image 2.0 Pro🇨🇳 CNFrontier

Alibaba

Img Elo1,236
Latency3.6s
$/1k$75/1k
Prompt %90.8%
vs GPT Image 2 (High)Specs
4Seedream 5.0 Pro🇨🇳 CNFrontier

ByteDance

Img Elo1,283
Latency3.4s
$/1k$90/1k
Prompt %92%
vs GPT Image 2 (High)Specs
#Model Image Elo Gen Time (s) $ / 1k imgs Prompt % Text % Compare vs
1
Wan2.6 Text to Image🇨🇳 CNOpen

Alibaba

1,2153.1s$30/1k89.5%87%
2
Seedream 4.0🇨🇳 CN

ByteDance

1,2252.8s$30/1k89%88%
3
Qwen Image 2.0 Pro🇨🇳 CNFrontier

Alibaba

1,2363.6s$75/1k90.8%92%
4
Seedream 5.0 Pro🇨🇳 CNFrontier

ByteDance

1,2833.4s$90/1k92%91.5%
4 models in this filter · sorted by image elo
All benchmark leaderboardsIntent rankings
US vs China AI FrontierGap Closed & Surpassing on Cost

How Close is China's AI to US Flagships?

Quality is virtually tied on independent benchmarks, while open-weight Chinese frontiers offer substantial price-to-performance advantages.

🇺🇸 US LeaderClaude Opus 5
vs
🇨🇳 China RivalGLM-5.3
General IntelligenceChatbot Arena human preference score
Claude Opus 51624 Elo
vs
GLM-5.31558 Elo
Near Parity (96%)
Software EngineeringSWE-bench real GitHub bug resolution
Claude Opus 579.2%
vs
GLM-5.376.4%
Near Parity (96%)
API Pricing & CostOutput price per 1 Million tokens
Claude Opus 5$25.00
vs
GLM-5.3$1.80
13.9× Cheaper
Model AvailabilityWeights download & self-hosting
Claude Opus 5Closed API
vs
GLM-5.3Open Weights
Open Source
Key Takeaway:GLM-5.3 matches 96% of Claude Opus 5's intelligence at 93% lower cost.
Full Showdown Fact Sheet

Leaders Matchup

Open GPT Image 2 (High) vs Midjourney v7 fact sheetOpen showdown fact sheet|View all comparisons
Multi-Dimensional Capability Radar

Frontier Trio Capability Radar

Comparing top Western standard models with China's leading frontier rival across 6 skill dimensions. Tap any spoke or dot to inspect.

Tap any node to inspect
Percentile 0–100
Verified Intelligence & Dispatches

Recent AI News & Benchmark Briefings

Independent analysis on model promotions, price reductions, and verified score movements.

View all dispatches (50)
Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama
compare
CompareLLM Intelligence DeskAug 16, 2026

Haiku 4.5 vs Llama 4 Scout: cheap Claude versus cheap open Llama

$1/$5 hosted Claude vs a cheap long-context open sibling. The alternative-to-Sonnet pair that is honest.

Shipped & Updated

New and Refreshed Models

View all 100 catalog models
Zhipu2026-08-14

GLM-5.3

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

Alibaba2026-08-14

Qwen3.8 27B

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

How to read this leaderboard

Plain-English methodology and leaderboard answers

Preference Elo is a crowd vote from LMArena / Arena. People see two hidden answers and pick the one they like more. The model that wins more often gets a higher Elo. That means people preferred it — not that it passed a school test. It is not SWE-bench, not accuracy, and not a number we invent.
Fastest
Opus 5 vs GPT-5.6
Open weights
Top Verified Frontiers
Top Preference Elo
1,624Claude Opus 5
Fastest TTFT
82 msGemini 3.7 Flash
Best Value (Elo ≥ 1450)
$0.14/1MYi-Lightning
Indexable Models
7163 Active

Head-to-Head Comparison Studio

All Comparisons Matrix
Trending Comparisons & RivalriesVerified Telemetry
Frontier Titan

Claude Opus 5 vs GPT-5.6 Sol

Top flagship duel for #1 general intelligence

East vs West

Claude Sonnet 5 vs DeepSeek V4

Western standard vs Chinese frontier price/perf leader

Open Weights

Qwen3.8 vs Llama 4 Scout

Top self-hostable open-weight models compared

High-Speed

GLM-5.3 vs Gemini 3.7 Flash

Fast reasoning & latency vs low $/1M API cost

Verified briefRead dispatch
Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE
launch
CompareLLM Intelligence DeskAug 14, 2026

Qwen3.8 27B (Aug 14): dense enough to self-host, not a frontier MoE

Alibaba open-weight 27B drop. Local/open-weight lists should see it. It will not win frontier-agents.

Verified briefRead dispatch
GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base
launch
CompareLLM Intelligence DeskAug 14, 2026

GLM-5.3 (Aug 14): Z.ai post-train of the 5.2 744B base

Coding-plan live; open weights promised after a two-week safety review. Newest Zhipu row.

Verified briefRead dispatch
Google2026-08-13

Gemini 3.7 Flash

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

xAI2026-08-12

Grok 4.6

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

ByteDance2026-08-10

Seed 2.1 Turbo

ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.

DeepSeek2026-07-31

DeepSeek V4 Flash

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.