CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Catalog Directory

AI Models Catalog

100 tracked models with dated benchmark snapshots, pairwise comparison hubs, and Stack Engine scoring.

Frontier Models OnlyOpen Weights OnlyInteractive Compare Hub
Showing 100 of 100 catalog models
Provider list prices are USD
1Claude Opus 5AnthropicFrontier

Anthropic current default flagship. Official API $5/$25 per 1M tokens and a 1M context window (Anthropic, Jul 24 2026).

Elo (Votes)1,624
SWE-bench (Code)79.2%
Input / 1M tokens$5.0
Output / 1M tokens$25
View Claude Opus 5 Specs & BenchmarksCompare vs All
2Claude Fable 5AnthropicFrontier

Anthropic top-tier X-High reasoning model engineered for heavy multi-step tasks. Official API $10/$50 per 1M tokens (Claude Platform pricing, Aug 2026).

Elo (Votes)1,616
SWE-bench (Code)80%
Input / 1M tokens$10
Output / 1M tokens$50
View Claude Fable 5 Specs & BenchmarksCompare vs All
3GPT-5.6 SolOpenAIFrontier

OpenAI 5.6 flagship tier. Official API $5/$30 per 1M tokens (OpenAI pricing, Jul 30 2026 update left Sol unchanged).

Elo (Votes)1,608
SWE-bench (Code)77.6%
Input / 1M tokens$5.0
Output / 1M tokens$30
View GPT-5.6 Sol Specs & BenchmarksCompare vs All
4Claude Opus 4.8AnthropicFrontier

Prior Opus generation still billed at $5/$25. Kept as a compare baseline against Opus 5.

Elo (Votes)1,598
SWE-bench (Code)78.4%
Input / 1M tokens$5.0
Output / 1M tokens$25
View Claude Opus 4.8 Specs & BenchmarksCompare vs All
5Grok 4.6xAIFrontier

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

Elo (Votes)1,592
SWE-bench (Code)69.1%
Input / 1M tokens$2.0
Output / 1M tokens$6.0
View Grok 4.6 Specs & BenchmarksCompare vs All
6Claude Opus 4.6AnthropicFrontier

Follow-on Opus release. Slightly behind 4.5 on official SWE-bench bash-only in the last published sweep; stronger multi-step autonomous agent reasoning.

Elo (Votes)1,574
SWE-bench (Code)75.9%
Input / 1M tokens$15
Output / 1M tokens$75
View Claude Opus 4.6 Specs & BenchmarksCompare vs All
7Gemini 3.6 ProGoogleFrontier

Current Google Pro-class multimodal model. Long context, strong coding, billed like the 3.x Pro tier.

Elo (Votes)1,570
SWE-bench (Code)74.6%
Input / 1M tokens$1.3
Output / 1M tokens$5.0
View Gemini 3.6 Pro Specs & BenchmarksCompare vs All
8Claude Opus 4.5AnthropicFrontier

Anthropic frontier coding and computer-use model. SWE-bench leader on the official mini-SWE-agent harness in the Feb 2026 refresh.

Elo (Votes)1,568
SWE-bench (Code)76.8%
Input / 1M tokens$15
Output / 1M tokens$75
View Claude Opus 4.5 Specs & BenchmarksCompare vs All
9OpenAI o3-miniOpenAIFrontier

OpenAI cost-efficient reasoning model with adjustable thinking effort and top coding benchmark scores.

Elo (Votes)1,560
SWE-bench (Code)78.5%
Input / 1M tokens$1.1
Output / 1M tokens$4.4
View OpenAI o3-mini Specs & BenchmarksCompare vs All
10GPT-5OpenAIFrontier

OpenAI flagship reasoning model for 2025–26. Strong general preference Elo and multimodal coverage.

Elo (Votes)1,558
SWE-bench (Code)68.4%
Input / 1M tokens$5.0
Output / 1M tokens$20
View GPT-5 Specs & BenchmarksCompare vs All
11GLM-5.3Zhipu🇨🇳 CNFrontier

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

Elo (Votes)1,558
SWE-bench (Code)76.4%
Input / 1M tokens$0.50
Output / 1M tokens$1.8
View GLM-5.3 Specs & BenchmarksCompare vs All
12Claude Sonnet 5AnthropicFrontier

Anthropic workhorse. Official $2/$10 per 1M tokens made permanent on Aug 10 2026.

Elo (Votes)1,556
SWE-bench (Code)73.8%
Input / 1M tokens$2.0
Output / 1M tokens$10
View Claude Sonnet 5 Specs & BenchmarksCompare vs All
13Gemini 3 ProGoogleFrontier

Google frontier multimodal model with a multi-million-token context window.

Elo (Votes)1,552
SWE-bench (Code)72.4%
Input / 1M tokens$1.3
Output / 1M tokens$5.0
View Gemini 3 Pro Specs & BenchmarksCompare vs All
14GPT-5.6 TerraOpenAIFrontier

OpenAI 5.6 mid tier. Official API $2/$12 per 1M after the Jul 30 2026 price cut.

Elo (Votes)1,548
SWE-bench (Code)70.4%
Input / 1M tokens$2.0
Output / 1M tokens$12
View GPT-5.6 Terra Specs & BenchmarksCompare vs All
15GPT-4.5 OrionOpenAIFrontier

OpenAI largest dense non-reasoning model with expansive world knowledge and reduced hallucinations.

Elo (Votes)1,540
SWE-bench (Code)71%
Input / 1M tokens$75
Output / 1M tokens$150
View GPT-4.5 Orion Specs & BenchmarksCompare vs All
16DeepSeek V4 ProDeepSeek🇨🇳 CNOpen WeightsFrontier

Open-weight-adjacent DeepSeek flagship. High reasoning density per dollar.

Elo (Votes)1,536
SWE-bench (Code)71.6%
Input / 1M tokens$0.40
Output / 1M tokens$1.6
View DeepSeek V4 Pro Specs & BenchmarksCompare vs All
17Gemini 3.7 FlashGoogleFrontier

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

Elo (Votes)1,530
SWE-bench (Code)72.4%
Input / 1M tokens$0.75
Output / 1M tokens$3.8
View Gemini 3.7 Flash Specs & BenchmarksCompare vs All
18Kimi K3Moonshot🇨🇳 CNFrontier

Moonshot AI frontier flagship reasoning and agent swarm model with 256k context and top-tier SWE-bench coding capability.

Elo (Votes)1,528
SWE-bench (Code)73.8%
Input / 1M tokens$0.60
Output / 1M tokens$2.5
View Kimi K3 Specs & BenchmarksCompare vs All
19Claude Sonnet 4.5AnthropicFrontier

Workhorse Anthropic model: most of Opus coding quality at a mid-tier price.

Elo (Votes)1,524
SWE-bench (Code)70.1%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Claude Sonnet 4.5 Specs & BenchmarksCompare vs All
20Qwen 3 MaxAlibaba🇨🇳 CNFrontier

Alibaba flagship. Strong math and multilingual code.

Elo (Votes)1,518
SWE-bench (Code)67.8%
Input / 1M tokens$1.2
Output / 1M tokens$4.8
View Qwen 3 Max Specs & BenchmarksCompare vs All
21Kimi K2.5 MaxMoonshot🇨🇳 CNFrontier

Moonshot flagship reasoning model with high-fidelity coding, deep thinking, and tool use across a 200k context.

Elo (Votes)1,515
SWE-bench (Code)71.3%
Input / 1M tokens$0.50
Output / 1M tokens$2.0
View Kimi K2.5 Max Specs & BenchmarksCompare vs All
22OpenAI o1OpenAI

OpenAI pioneer reasoning foundation model designed for complex science, math, and multi-step reasoning.

Elo (Votes)1,512
SWE-bench (Code)48.9%
Input / 1M tokens$15
Output / 1M tokens$60
View OpenAI o1 Specs & BenchmarksCompare vs All
23Gemini 3.6 FlashGoogleFrontier

Google Jul 21 workhorse. Now shares the 3.7 Flash introductory $0.75/$3.75 rate through Dec 31 2026.

Elo (Votes)1,506
SWE-bench (Code)70.8%
Input / 1M tokens$0.75
Output / 1M tokens$3.8
View Gemini 3.6 Flash Specs & BenchmarksCompare vs All
24Gemini 3 FlashGoogleFrontier

Fast Google frontier-adjacent model. Near-top official SWE-bench at a fraction of Opus price.

Elo (Votes)1,495
SWE-bench (Code)75.8%
Input / 1M tokens$0.15
Output / 1M tokens$0.60
View Gemini 3 Flash Specs & BenchmarksCompare vs All
25Qwen QwQ 32BAlibaba🇨🇳 CNOpen WeightsFrontier

Alibaba specialized open reasoning model competing with frontier closed reasoning models.

Elo (Votes)1,495
SWE-bench (Code)67.5%
Input / 1M tokens$0.30
Output / 1M tokens$1.2
View Qwen QwQ 32B Specs & BenchmarksCompare vs All
26GLM-5.2Zhipu🇨🇳 CNOpen WeightsFrontier

Zhipu flagship. Strong Chinese/English coding and agents.

Elo (Votes)1,492
SWE-bench (Code)69.3%
Input / 1M tokens$0.50
Output / 1M tokens$1.8
View GLM-5.2 Specs & BenchmarksCompare vs All
27Gemini 2.0 Flash ThinkingGoogleFrontier

Google experimental reasoning model that visualizes thoughts in real-time.

Elo (Votes)1,490
SWE-bench (Code)66.8%
Input / 1M tokens$0.10
Output / 1M tokens$0.40
View Gemini 2.0 Flash Thinking Specs & BenchmarksCompare vs All
28DeepSeek V4 FlashDeepSeek🇨🇳 CNOpen WeightsFrontier

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.

Elo (Votes)1,490
SWE-bench (Code)71.2%
Input / 1M tokens$0.14
Output / 1M tokens$0.28
View DeepSeek V4 Flash Specs & BenchmarksCompare vs All
29Claude Opus 4Anthropic

First Claude 4 Opus generation. Baseline for opus 4 vs opus 5.

Elo (Votes)1,490
SWE-bench (Code)72.5%
Input / 1M tokens$15
Output / 1M tokens$75
View Claude Opus 4 Specs & BenchmarksCompare vs All
30Grok 4xAI

Previous xAI flagship.

Elo (Votes)1,488
SWE-bench (Code)58.6%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Grok 4 Specs & BenchmarksCompare vs All
31MiniMax M2.5MiniMax🇨🇳 CNOpen WeightsFrontier

MiniMax coding model. Tied near the top of official SWE-bench bash-only in Feb 2026.

Elo (Votes)1,478
SWE-bench (Code)75.8%
Input / 1M tokens$0.30
Output / 1M tokens$1.2
View MiniMax M2.5 Specs & BenchmarksCompare vs All
32Yi-Lightning01.AI🇨🇳 CNFrontier

01.AI ultra-fast reasoning model delivering top LiveBench efficiency.

Elo (Votes)1,475
SWE-bench (Code)63.4%
Input / 1M tokens$0.14
Output / 1M tokens$0.14
View Yi-Lightning Specs & BenchmarksCompare vs All
33GLM-5V TurboZhipu🇨🇳 CNOpen WeightsFrontier

Zhipu high-throughput multimodal vision model with sub-second latency and competitive coding.

Elo (Votes)1,475
SWE-bench (Code)65.2%
Input / 1M tokens$0.15
Output / 1M tokens$0.60
View GLM-5V Turbo Specs & BenchmarksCompare vs All
34Seed 2.1 TurboByteDance🇨🇳 CN

ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.

Elo (Votes)1,472
SWE-bench (Code)63.2%
Input / 1M tokens$0.27
Output / 1M tokens$1.1
View Seed 2.1 Turbo Specs & BenchmarksCompare vs All
35Doubao Pro 1.5ByteDance🇨🇳 CNFrontier

ByteDance flagship enterprise model with ultra-low token cost and 128k context.

Elo (Votes)1,470
SWE-bench (Code)62.1%
Input / 1M tokens$0.11
Output / 1M tokens$0.28
View Doubao Pro 1.5 Specs & BenchmarksCompare vs All
36Llama 4 MaverickMetaOpen WeightsFrontier

Meta natively multimodal open-weight flagship.

Elo (Votes)1,468
SWE-bench (Code)57.1%
Input / 1M tokens$0.20
Output / 1M tokens$0.60
View Llama 4 Maverick Specs & BenchmarksCompare vs All
37GPT-5.6 LunaOpenAI

OpenAI 5.6 fast/cheap tier. Official API $0.20/$1.20 per 1M after the Jul 30 80% Luna cut.

Elo (Votes)1,466
SWE-bench (Code)61.8%
Input / 1M tokens$0.20
Output / 1M tokens$1.2
View GPT-5.6 Luna Specs & BenchmarksCompare vs All
38Kimi K2Moonshot🇨🇳 CN

Moonshot foundation model with strong native bilingual reasoning and 65.8% SWE-bench Verified coding baseline.

Elo (Votes)1,465
SWE-bench (Code)65.8%
Input / 1M tokens$0.35
Output / 1M tokens$1.4
View Kimi K2 Specs & BenchmarksCompare vs All
39Mistral Large 3MistralFrontier

Mistral European flagship with strong function calling.

Elo (Votes)1,456
SWE-bench (Code)55.4%
Input / 1M tokens$2.0
Output / 1M tokens$6.0
View Mistral Large 3 Specs & BenchmarksCompare vs All
40Claude Sonnet 4Anthropic

First Sonnet 4 generation. Bridge between 3.5/3.7 and Sonnet 5.

Elo (Votes)1,455
SWE-bench (Code)62.8%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Claude Sonnet 4 Specs & BenchmarksCompare vs All
41Ernie 4.5 TurboBaidu🇨🇳 CN

Baidu current multimodal enterprise foundation model with broad Chinese knowledge.

Elo (Votes)1,450
SWE-bench (Code)57%
Input / 1M tokens$0.30
Output / 1M tokens$1.2
View Ernie 4.5 Turbo Specs & BenchmarksCompare vs All
42Qwen 2.5 PlusAlibaba🇨🇳 CN

Alibaba balanced flagship API model with high-throughput general reasoning.

Elo (Votes)1,445
SWE-bench (Code)59.2%
Input / 1M tokens$0.25
Output / 1M tokens$1.0
View Qwen 2.5 Plus Specs & BenchmarksCompare vs All
43OpenAI o1-miniOpenAI

OpenAI high-speed, cost-effective reasoning model optimized for STEM, math, and code generation.

Elo (Votes)1,445
SWE-bench (Code)56.4%
Input / 1M tokens$1.1
Output / 1M tokens$4.4
View OpenAI o1-mini Specs & BenchmarksCompare vs All
44Qwen3 235BAlibaba🇨🇳 CNOpen Weights

Open-weight Qwen3 mixture-of-experts.

Elo (Votes)1,440
SWE-bench (Code)59.4%
Input / 1M tokens$0.22
Output / 1M tokens$0.88
View Qwen3 235B Specs & BenchmarksCompare vs All
45Claude Haiku 4.5Anthropic

Anthropic cheap/fast Claude SKU. Official Claude Platform list price $1/$5 per 1M tokens (Anthropic Haiku page).

Elo (Votes)1,438
SWE-bench (Code)58.2%
Input / 1M tokens$1.0
Output / 1M tokens$5.0
View Claude Haiku 4.5 Specs & BenchmarksCompare vs All
46Yi-Large01.AI🇨🇳 CN

01.AI full-scale dense model for complex instruction following.

Elo (Votes)1,430
SWE-bench (Code)53.8%
Input / 1M tokens$0.30
Output / 1M tokens$0.30
View Yi-Large Specs & BenchmarksCompare vs All
47Qwen 2.5 Coder 32BAlibaba🇨🇳 CNOpen Weights

Alibaba dedicated open-weight code generation model with near-frontier SWE-bench Verified coding capability.

Elo (Votes)1,425
SWE-bench (Code)65.2%
Input / 1M tokens$0.18
Output / 1M tokens$0.72
View Qwen 2.5 Coder 32B Specs & BenchmarksCompare vs All
48Kimi Chat 1.5Moonshot🇨🇳 CN

Moonshot ultra-long context model supporting up to 2 million tokens per request.

Elo (Votes)1,420
SWE-bench (Code)51%
Input / 1M tokens$0.20
Output / 1M tokens$0.80
View Kimi Chat 1.5 Specs & BenchmarksCompare vs All
49Llama 4 ScoutMetaOpen Weights

Meta open-weight Llama 4 long-context sibling of Maverick. Common public API lists sit near $0.08–$0.30 / $0.30–$0.70 per 1M; we store a conservative hosted list until OpenRouter overwrites.

Elo (Votes)1,412
SWE-bench (Code)52.4%
Input / 1M tokens$0.080
Output / 1M tokens$0.30
View Llama 4 Scout Specs & BenchmarksCompare vs All
50GPT-5 miniOpenAI

Cost-efficient GPT-5 distill for high-volume agents.

Elo (Votes)1,410
SWE-bench (Code)52.4%
Input / 1M tokens$0.25
Output / 1M tokens$2.0
View GPT-5 mini Specs & BenchmarksCompare vs All
51Grok 3xAI

Previous xAI generation before Grok 4. Kept for grok 3 vs grok 4.

Elo (Votes)1,405
SWE-bench (Code)51.4%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Grok 3 Specs & BenchmarksCompare vs All
52Qwen3.8 27BAlibaba🇨🇳 CNOpen Weights

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

Elo (Votes)1,398
SWE-bench (Code)58.8%
Input / 1M tokens$0.10
Output / 1M tokens$0.40
View Qwen3.8 27B Specs & BenchmarksCompare vs All
53Llama 3.1 405BMetaOpen Weights

Meta flagship open-weight 405B dense foundation model with 128k context window.

Elo (Votes)1,370
SWE-bench (Code)58.4%
Input / 1M tokens$1.3
Output / 1M tokens$3.5
View Llama 3.1 405B Specs & BenchmarksCompare vs All
54DeepSeek Coder V2DeepSeek🇨🇳 CNOpen Weights

DeepSeek open-weight Mixture-of-Experts coding model supporting 338 programming languages and 128k context.

Elo (Votes)1,365
SWE-bench (Code)60.5%
Input / 1M tokens$0.14
Output / 1M tokens$0.28
View DeepSeek Coder V2 Specs & BenchmarksCompare vs All
55Claude 3.7 SonnetAnthropic

Hybrid-reasoning Sonnet from 2025. Kept for historical compare pages.

Elo (Votes)1,362
SWE-bench (Code)70.3%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Claude 3.7 Sonnet Specs & BenchmarksCompare vs All
56Doubao Lite 1.5ByteDance🇨🇳 CN

ByteDance high-speed lightweight model priced at sub-cent levels.

Elo (Votes)1,360
SWE-bench (Code)46%
Input / 1M tokens$0.040
Output / 1M tokens$0.10
View Doubao Lite 1.5 Specs & BenchmarksCompare vs All
57DeepSeek R1DeepSeek🇨🇳 CNOpen Weights

Open-weights reasoning model trained with large-scale RL.

Elo (Votes)1,358
SWE-bench (Code)65.2%
Input / 1M tokens$0.55
Output / 1M tokens$2.2
View DeepSeek R1 Specs & BenchmarksCompare vs All
58Gemini 2.5 ProGoogle

Previous Google long-context flagship.

Elo (Votes)1,350
SWE-bench (Code)63.8%
Input / 1M tokens$1.3
Output / 1M tokens$5.0
View Gemini 2.5 Pro Specs & BenchmarksCompare vs All
59Command ACohere

Cohere enterprise RAG and tool-use model.

Elo (Votes)1,344
SWE-bench (Code)42.6%
Input / 1M tokens$2.5
Output / 1M tokens$10
View Command A Specs & BenchmarksCompare vs All
60GPT-4oOpenAI

Previous OpenAI flagship. Still a common compare baseline on legacy pages.

Elo (Votes)1,335
SWE-bench (Code)54.8%
Input / 1M tokens$2.5
Output / 1M tokens$10
View GPT-4o Specs & BenchmarksCompare vs All
61Codestral 25.01Mistral

Mistral code-specialist model.

Elo (Votes)1,320
SWE-bench (Code)51.8%
Input / 1M tokens$0.30
Output / 1M tokens$0.90
View Codestral 25.01 Specs & BenchmarksCompare vs All
62DeepSeek V3DeepSeek🇨🇳 CNOpen Weights

Prior DeepSeek flagship. Baseline for v3 vs v4.

Elo (Votes)1,310
SWE-bench (Code)48.6%
Input / 1M tokens$0.27
Output / 1M tokens$1.1
View DeepSeek V3 Specs & BenchmarksCompare vs All
63Gemini 2.5 FlashGoogle

Previous Google speed workhorse.

Elo (Votes)1,295
SWE-bench (Code)48%
Input / 1M tokens$0.10
Output / 1M tokens$0.40
View Gemini 2.5 Flash Specs & BenchmarksCompare vs All
64Llama 3.3 70BMetaOpen Weights

Previous Meta 70B open-weight workhorse.

Elo (Votes)1,285
SWE-bench (Code)45.1%
Input / 1M tokens$0.18
Output / 1M tokens$0.59
View Llama 3.3 70B Specs & BenchmarksCompare vs All
65Claude 3.5 SonnetAnthropic

2024 workhorse. Still a high-intent compare against GPT-4o.

Elo (Votes)1,280
SWE-bench (Code)49%
Input / 1M tokens$3.0
Output / 1M tokens$15
View Claude 3.5 Sonnet Specs & BenchmarksCompare vs All
66GPT-4o miniOpenAI

Legacy small OpenAI model. Useful as a cheap baseline.

Elo (Votes)1,272
SWE-bench (Code)41.2%
Input / 1M tokens$0.15
Output / 1M tokens$0.60
View GPT-4o mini Specs & BenchmarksCompare vs All
67Gemini 1.5 ProGoogle

First million-token Gemini Pro. Baseline for 1.5 vs 2.5 vs 3.x Pro.

Elo (Votes)1,260
SWE-bench (Code)38%
Input / 1M tokens$1.3
Output / 1M tokens$5.0
View Gemini 1.5 Pro Specs & BenchmarksCompare vs All
68GPT-4 TurboOpenAI

GPT-4 Turbo 128k. Historical flagship for gpt-4 turbo vs gpt-4o / gpt-5.

Elo (Votes)1,255
SWE-bench (Code)33.2%
Input / 1M tokens$10
Output / 1M tokens$30
View GPT-4 Turbo Specs & BenchmarksCompare vs All
69Claude 3 OpusAnthropic

Original Claude 3 flagship. Kept so opus 3 vs later Opus and vs GPT-4o still resolve.

Elo (Votes)1,248
SWE-bench (Code)38.4%
Input / 1M tokens$15
Output / 1M tokens$75
View Claude 3 Opus Specs & BenchmarksCompare vs All
70Llama 3.1 70BMetaOpen Weights

Llama 3.1 70B instruct. Predecessor to 3.3 70B and Llama 4.

Elo (Votes)1,240
SWE-bench (Code)40.2%
Input / 1M tokens$0.35
Output / 1M tokens$0.40
View Llama 3.1 70B Specs & BenchmarksCompare vs All
71Claude 3.5 HaikuAnthropic

Previous cheap Claude. Haiku 4.5 is the current $1/$5 SKU.

Elo (Votes)1,224
SWE-bench (Code)36.2%
Input / 1M tokens$0.80
Output / 1M tokens$4.0
View Claude 3.5 Haiku Specs & BenchmarksCompare vs All
72GPT Image 2 (High)OpenAIFrontier

OpenAI flagship diffusion-transformer image synthesis model with supreme prompt adherence, realistic textures, and complex text composition.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View GPT Image 2 (High) Specs & BenchmarksCompare vs All
73GPT Image 2 (Low / Fast)OpenAI

Fast distilled tier of GPT Image 2 designed for high-throughput interactive creative workflows at 70% lower price.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View GPT Image 2 (Low / Fast) Specs & BenchmarksCompare vs All
74GPT Image 1.5OpenAI

Previous OpenAI image generation flagship. High fidelity with proven enterprise reliability.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View GPT Image 1.5 Specs & BenchmarksCompare vs All
75DALL-E 3OpenAI

Legacy OpenAI image model benchmark baseline. Retained for historical comparisons.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View DALL-E 3 Specs & BenchmarksCompare vs All
76Reve 2.1ReveFrontier

State-of-the-art cinematic image synthesis foundation model known for photorealistic lighting and aesthetic composition.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Reve 2.1 Specs & BenchmarksCompare vs All
77Nano Banana Pro (Gemini 3 Pro Image)GoogleFrontier

Google DeepMind frontier multimodal image generation flagship with Deep Research and complex multi-object spatial reasoning.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Nano Banana Pro (Gemini 3 Pro Image) Specs & BenchmarksCompare vs All
78Nano Banana 2 (Gemini 3.1 Flash Image)GoogleFrontier

Google high-speed multimodal generative model delivering top Elo performance at half the latency and cost of Pro.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Nano Banana 2 (Gemini 3.1 Flash Image) Specs & BenchmarksCompare vs All
79Nano Banana 2 LiteGoogle

Sub-1.5 second ultra-low latency tier for real-time applications and mobile game asset pipelines.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Nano Banana 2 Lite Specs & BenchmarksCompare vs All
80FLUX.2 [max]Black Forest LabsFrontier

Black Forest Labs maximum-capacity flow-matching model with photoreal anatomy and leading typography fidelity.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View FLUX.2 [max] Specs & BenchmarksCompare vs All
81FLUX.2 [flex]Black Forest LabsFrontier

Flexible step-distilled variant of FLUX.2 offering 99% of Max quality with 45% faster generation speeds.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View FLUX.2 [flex] Specs & BenchmarksCompare vs All
82FLUX.1.1 [pro]Black Forest Labs

First-generation professional API model from Black Forest Labs, renowned for 6x faster generation than original FLUX.1 Pro.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View FLUX.1.1 [pro] Specs & BenchmarksCompare vs All
83FLUX.1 [dev]Black Forest LabsOpen Weights

Open-weights non-commercial base model with 12B parameters, serving as the foundation for the open-source fine-tuning ecosystem.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View FLUX.1 [dev] Specs & BenchmarksCompare vs All
84FLUX.1 [schnell]Black Forest LabsOpen Weights

Apache 2.0 4-step distilled open weights model optimized for local inference and instant preview generation.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View FLUX.1 [schnell] Specs & BenchmarksCompare vs All
85Ideogram 4.0 (Quality)IdeogramFrontier

Industry-leading typography, graphic design, and in-image text layout model with precise kerning and complex banner generation.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Ideogram 4.0 (Quality) Specs & BenchmarksCompare vs All
86Ideogram 4.0 Open WeightsIdeogramOpen WeightsFrontier

Open-weight release of Ideogram 4.0, bringing world-class text rendering to self-hosted enterprise infrastructure.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Ideogram 4.0 Open Weights Specs & BenchmarksCompare vs All
87Recraft V4.1 Utility ProRecraftFrontier

Specialized generative model for vector art, icon sets, illustrations, and commercial design assets with native brand color matching.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Recraft V4.1 Utility Pro Specs & BenchmarksCompare vs All
88Recraft V4.1 UtilityRecraft

High-speed utility tier of Recraft V4.1 tailored for rapid asset production at $0.035/image.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Recraft V4.1 Utility Specs & BenchmarksCompare vs All
89Krea 2 LargeKreaFrontier

High-resolution real-time generation model tuned for artistic composition and dynamic concept art.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Krea 2 Large Specs & BenchmarksCompare vs All
90Krea 2 Medium TurboKrea

Sub-2 second interactive generation engine delivering ultra-affordable $15/1k image generation.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Krea 2 Medium Turbo Specs & BenchmarksCompare vs All
91Qwen Image 2.0 ProAlibaba🇨🇳 CNFrontier

Alibaba multimodal image synthesis foundation model with strong bilingual Chinese/English typography and cultural asset fidelity.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Qwen Image 2.0 Pro Specs & BenchmarksCompare vs All
92Wan2.6 Text to ImageAlibaba🇨🇳 CNOpen Weights

Alibaba high-efficiency open video/image unified transformer backbone.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Wan2.6 Text to Image Specs & BenchmarksCompare vs All
93Seedream 5.0 ProByteDance🇨🇳 CNFrontier

ByteDance flagship image generation model with high photorealism and fine facial structure rendering.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Seedream 5.0 Pro Specs & BenchmarksCompare vs All
94Seedream 4.0ByteDance🇨🇳 CN

High-value ByteDance image foundation model at $30/1k images.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Seedream 4.0 Specs & BenchmarksCompare vs All
95Midjourney v7MidjourneyFrontier

Midjourney v7 premier creative image generation model with unparalleled stylistic nuances, coherent hands/anatomy, and cinematic grading.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Midjourney v7 Specs & BenchmarksCompare vs All
96Midjourney v6.1Midjourney

Previous standard in aesthetic digital art generation, retained as a historical compare baseline.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Midjourney v6.1 Specs & BenchmarksCompare vs All
97MAI-Image-2.5Microsoft AIFrontier

Microsoft AI proprietary foundation image model for Copilot Studio and enterprise creative workflows.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View MAI-Image-2.5 Specs & BenchmarksCompare vs All
98MAI-Image-2.5-FlashMicrosoft AI

Cost-optimized Microsoft AI image model delivering sub-2 second responses at $20/1k images.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View MAI-Image-2.5-Flash Specs & BenchmarksCompare vs All
99HiDream-O1-Image-1.5HiDream

HiDream high-definition visual generator optimized for photorealism and accurate complex lighting.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View HiDream-O1-Image-1.5 Specs & BenchmarksCompare vs All
100Luma UNI 1 MaxLuma LabsFrontier

Luma Labs universal 3D-aware image synthesis engine with spatial geometry consistency.

Elo (Votes)—
SWE-bench (Code)—
Input / 1M tokens—
Output / 1M tokens—
View Luma UNI 1 Max Specs & BenchmarksCompare vs All