CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Top Tier Capability

Frontier AI Models

48 state-of-the-art models actively pushing the frontier in reasoning, coding, and autonomous benchmark scores.

#1Claude Opus 5Anthropic
🏆 #1 Frontier Leader

Anthropic current default flagship. Official API $5/$25 per 1M tokens and a 1M context window (Anthropic, Jul 24 2026).

Elo:1,624
SWE-bench:79.2%
Out:$25/1M
#2Claude Fable 5Anthropic
vs #1 Claude Opus 5

Anthropic top-tier X-High reasoning model engineered for heavy multi-step tasks. Official API $10/$50 per 1M tokens (Claude Platform pricing, Aug 2026).

Elo:1,616
SWE-bench:80%
Out:$50/1M
#3GPT-5.6 SolOpenAI
vs #1 Claude Opus 5

OpenAI 5.6 flagship tier. Official API $5/$30 per 1M tokens (OpenAI pricing, Jul 30 2026 update left Sol unchanged).

Elo:1,608
SWE-bench:77.6%
Out:$30/1M
#4Claude Opus 4.8Anthropic
vs #1 Claude Opus 5

Prior Opus generation still billed at $5/$25. Kept as a compare baseline against Opus 5.

Elo:1,598
SWE-bench:78.4%
Out:$25/1M
#5Grok 4.6xAI
vs #1 Claude Opus 5

xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.

Elo:1,592
SWE-bench:69.1%
Out:$6/1M
#6Claude Opus 4.6Anthropic
vs #1 Claude Opus 5

Follow-on Opus release. Slightly behind 4.5 on official SWE-bench bash-only in the last published sweep; stronger multi-step autonomous agent reasoning.

Elo:1,574
SWE-bench:75.9%
Out:$75/1M
#7Gemini 3.6 ProGoogle
vs #1 Claude Opus 5

Current Google Pro-class multimodal model. Long context, strong coding, billed like the 3.x Pro tier.

Elo:1,570
SWE-bench:74.6%
Out:$5/1M
#8Claude Opus 4.5Anthropic
vs #1 Claude Opus 5

Anthropic frontier coding and computer-use model. SWE-bench leader on the official mini-SWE-agent harness in the Feb 2026 refresh.

Elo:1,568
SWE-bench:76.8%
Out:$75/1M
#9OpenAI o3-miniOpenAI
vs #1 Claude Opus 5

OpenAI cost-efficient reasoning model with adjustable thinking effort and top coding benchmark scores.

Elo:1,560
SWE-bench:78.5%
Out:$4.4/1M
#10GPT-5OpenAI
vs #1 Claude Opus 5

OpenAI flagship reasoning model for 2025–26. Strong general preference Elo and multimodal coverage.

Elo:1,558
SWE-bench:68.4%
Out:$20/1M
#11GLM-5.3Zhipu
vs #1 Claude Opus 5

Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).

Elo:1,558
SWE-bench:76.4%
Out:$1.8/1M
#12Claude Sonnet 5Anthropic
vs #1 Claude Opus 5

Anthropic workhorse. Official $2/$10 per 1M tokens made permanent on Aug 10 2026.

Elo:1,556
SWE-bench:73.8%
Out:$10/1M
#13Gemini 3 ProGoogle
vs #1 Claude Opus 5

Google frontier multimodal model with a multi-million-token context window.

Elo:1,552
SWE-bench:72.4%
Out:$5/1M
#14GPT-5.6 TerraOpenAI
vs #1 Claude Opus 5

OpenAI 5.6 mid tier. Official API $2/$12 per 1M after the Jul 30 2026 price cut.

Elo:1,548
SWE-bench:70.4%
Out:$12/1M
#15GPT-4.5 OrionOpenAI
vs #1 Claude Opus 5

OpenAI largest dense non-reasoning model with expansive world knowledge and reduced hallucinations.

Elo:1,540
SWE-bench:71%
Out:$150/1M
#16DeepSeek V4 ProDeepSeek
vs #1 Claude Opus 5

Open-weight-adjacent DeepSeek flagship. High reasoning density per dollar.

Elo:1,536
SWE-bench:71.6%
Out:$1.6/1M
#17Gemini 3.7 FlashGoogle
vs #1 Claude Opus 5

Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).

Elo:1,530
SWE-bench:72.4%
Out:$3.75/1M
#18Kimi K3Moonshot
vs #1 Claude Opus 5

Moonshot AI frontier flagship reasoning and agent swarm model with 256k context and top-tier SWE-bench coding capability.

Elo:1,528
SWE-bench:73.8%
Out:$2.5/1M
#19Claude Sonnet 4.5Anthropic
vs #1 Claude Opus 5

Workhorse Anthropic model: most of Opus coding quality at a mid-tier price.

Elo:1,524
SWE-bench:70.1%
Out:$15/1M
#20Qwen 3 MaxAlibaba
vs #1 Claude Opus 5

Alibaba flagship. Strong math and multilingual code.

Elo:1,518
SWE-bench:67.8%
Out:$4.8/1M
#21Kimi K2.5 MaxMoonshot
vs #1 Claude Opus 5

Moonshot flagship reasoning model with high-fidelity coding, deep thinking, and tool use across a 200k context.

Elo:1,515
SWE-bench:71.3%
Out:$2/1M
#22Gemini 3.6 FlashGoogle
vs #1 Claude Opus 5

Google Jul 21 workhorse. Now shares the 3.7 Flash introductory $0.75/$3.75 rate through Dec 31 2026.

Elo:1,506
SWE-bench:70.8%
Out:$3.75/1M
#23Gemini 3 FlashGoogle
vs #1 Claude Opus 5

Fast Google frontier-adjacent model. Near-top official SWE-bench at a fraction of Opus price.

Elo:1,495
SWE-bench:75.8%
Out:$0.6/1M
#24Qwen QwQ 32BAlibaba
vs #1 Claude Opus 5

Alibaba specialized open reasoning model competing with frontier closed reasoning models.

Elo:1,495
SWE-bench:67.5%
Out:$1.2/1M
#25GLM-5.2Zhipu
vs #1 Claude Opus 5

Zhipu flagship. Strong Chinese/English coding and agents.

Elo:1,492
SWE-bench:69.3%
Out:$1.8/1M
#26Gemini 2.0 Flash ThinkingGoogle
vs #1 Claude Opus 5

Google experimental reasoning model that visualizes thoughts in real-time.

Elo:1,490
SWE-bench:66.8%
Out:$0.4/1M
#27DeepSeek V4 FlashDeepSeek
vs #1 Claude Opus 5

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.

Elo:1,490
SWE-bench:71.2%
Out:$0.28/1M
#28MiniMax M2.5MiniMax
vs #1 Claude Opus 5

MiniMax coding model. Tied near the top of official SWE-bench bash-only in Feb 2026.

Elo:1,478
SWE-bench:75.8%
Out:$1.2/1M
#29Yi-Lightning01.AI
vs #1 Claude Opus 5

01.AI ultra-fast reasoning model delivering top LiveBench efficiency.

Elo:1,475
SWE-bench:63.4%
Out:$0.14/1M
#30GLM-5V TurboZhipu
vs #1 Claude Opus 5

Zhipu high-throughput multimodal vision model with sub-second latency and competitive coding.

Elo:1,475
SWE-bench:65.2%
Out:$0.6/1M
#31Doubao Pro 1.5ByteDance
vs #1 Claude Opus 5

ByteDance flagship enterprise model with ultra-low token cost and 128k context.

Elo:1,470
SWE-bench:62.1%
Out:$0.28/1M
#32Llama 4 MaverickMeta
vs #1 Claude Opus 5

Meta natively multimodal open-weight flagship.

Elo:1,468
SWE-bench:57.1%
Out:$0.6/1M
#33Mistral Large 3Mistral
vs #1 Claude Opus 5

Mistral European flagship with strong function calling.

Elo:1,456
SWE-bench:55.4%
Out:$6/1M
#34GPT Image 2 (High)OpenAI
vs #1 Claude Opus 5

OpenAI flagship diffusion-transformer image synthesis model with supreme prompt adherence, realistic textures, and complex text composition.

#35Reve 2.1Reve
vs #1 Claude Opus 5

State-of-the-art cinematic image synthesis foundation model known for photorealistic lighting and aesthetic composition.

#36Nano Banana Pro (Gemini 3 Pro Image)Google
vs #1 Claude Opus 5

Google DeepMind frontier multimodal image generation flagship with Deep Research and complex multi-object spatial reasoning.

#37Nano Banana 2 (Gemini 3.1 Flash Image)Google
vs #1 Claude Opus 5

Google high-speed multimodal generative model delivering top Elo performance at half the latency and cost of Pro.

#38FLUX.2 [max]Black Forest Labs
vs #1 Claude Opus 5

Black Forest Labs maximum-capacity flow-matching model with photoreal anatomy and leading typography fidelity.

#39FLUX.2 [flex]Black Forest Labs
vs #1 Claude Opus 5

Flexible step-distilled variant of FLUX.2 offering 99% of Max quality with 45% faster generation speeds.

#40Ideogram 4.0 (Quality)Ideogram
vs #1 Claude Opus 5

Industry-leading typography, graphic design, and in-image text layout model with precise kerning and complex banner generation.

#41Ideogram 4.0 Open WeightsIdeogram
vs #1 Claude Opus 5

Open-weight release of Ideogram 4.0, bringing world-class text rendering to self-hosted enterprise infrastructure.

#42Recraft V4.1 Utility ProRecraft
vs #1 Claude Opus 5

Specialized generative model for vector art, icon sets, illustrations, and commercial design assets with native brand color matching.

#43Krea 2 LargeKrea
vs #1 Claude Opus 5

High-resolution real-time generation model tuned for artistic composition and dynamic concept art.

#44Qwen Image 2.0 ProAlibaba
vs #1 Claude Opus 5

Alibaba multimodal image synthesis foundation model with strong bilingual Chinese/English typography and cultural asset fidelity.

#45Seedream 5.0 ProByteDance
vs #1 Claude Opus 5

ByteDance flagship image generation model with high photorealism and fine facial structure rendering.

#46Midjourney v7Midjourney
vs #1 Claude Opus 5

Midjourney v7 premier creative image generation model with unparalleled stylistic nuances, coherent hands/anatomy, and cinematic grading.

#47MAI-Image-2.5Microsoft AI
vs #1 Claude Opus 5

Microsoft AI proprietary foundation image model for Copilot Studio and enterprise creative workflows.

#48Luma UNI 1 MaxLuma Labs
vs #1 Claude Opus 5

Luma Labs universal 3D-aware image synthesis engine with spatial geometry consistency.