CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
Open Weights & Self-Hostable

Open-Weight AI Models

21 open-weights models ranked by preference Elo and SWE-bench. Read our open vs closed evaluation guide.

#1DeepSeek V4 ProDeepSeek
Compare vs all

Open-weight-adjacent DeepSeek flagship. High reasoning density per dollar.

Elo:1,536
SWE-bench:71.6%
Out:$1.6/1M
#2Qwen QwQ 32BAlibaba
Compare vs all

Alibaba specialized open reasoning model competing with frontier closed reasoning models.

Elo:1,495
SWE-bench:67.5%
Out:$1.2/1M
#3GLM-5.2Zhipu
Compare vs all

Zhipu flagship. Strong Chinese/English coding and agents.

Elo:1,492
SWE-bench:69.3%
Out:$1.8/1M
#4DeepSeek V4 FlashDeepSeek
Compare vs all

DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.

Elo:1,490
SWE-bench:71.2%
Out:$0.28/1M
#5MiniMax M2.5MiniMax
Compare vs all

MiniMax coding model. Tied near the top of official SWE-bench bash-only in Feb 2026.

Elo:1,478
SWE-bench:75.8%
Out:$1.2/1M
#6GLM-5V TurboZhipu
Compare vs all

Zhipu high-throughput multimodal vision model with sub-second latency and competitive coding.

Elo:1,475
SWE-bench:65.2%
Out:$0.6/1M
#7Llama 4 MaverickMeta
Compare vs all

Meta natively multimodal open-weight flagship.

Elo:1,468
SWE-bench:57.1%
Out:$0.6/1M
#8Qwen3 235BAlibaba
Compare vs all

Open-weight Qwen3 mixture-of-experts.

Elo:1,440
SWE-bench:59.4%
Out:$0.88/1M
#9Qwen 2.5 Coder 32BAlibaba
Compare vs all

Alibaba dedicated open-weight code generation model with near-frontier SWE-bench Verified coding capability.

Elo:1,425
SWE-bench:65.2%
Out:$0.72/1M
#10Llama 4 ScoutMeta
Compare vs all

Meta open-weight Llama 4 long-context sibling of Maverick. Common public API lists sit near $0.08–$0.30 / $0.30–$0.70 per 1M; we store a conservative hosted list until OpenRouter overwrites.

Elo:1,412
SWE-bench:52.4%
Out:$0.3/1M
#11Qwen3.8 27BAlibaba
Compare vs all

Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.

Elo:1,398
SWE-bench:58.8%
Out:$0.4/1M
#12Llama 3.1 405BMeta
Compare vs all

Meta flagship open-weight 405B dense foundation model with 128k context window.

Elo:1,370
SWE-bench:58.4%
Out:$3.5/1M
#13DeepSeek Coder V2DeepSeek
Compare vs all

DeepSeek open-weight Mixture-of-Experts coding model supporting 338 programming languages and 128k context.

Elo:1,365
SWE-bench:60.5%
Out:$0.28/1M
#14DeepSeek R1DeepSeek
Compare vs all

Open-weights reasoning model trained with large-scale RL.

Elo:1,358
SWE-bench:65.2%
Out:$2.19/1M
#15DeepSeek V3DeepSeek
Compare vs all

Prior DeepSeek flagship. Baseline for v3 vs v4.

Elo:1,310
SWE-bench:48.6%
Out:$1.1/1M
#16Llama 3.3 70BMeta
Compare vs all

Previous Meta 70B open-weight workhorse.

Elo:1,285
SWE-bench:45.1%
Out:$0.59/1M
#17Llama 3.1 70BMeta
Compare vs all

Llama 3.1 70B instruct. Predecessor to 3.3 70B and Llama 4.

Elo:1,240
SWE-bench:40.2%
Out:$0.4/1M
#18FLUX.1 [dev]Black Forest Labs
Compare vs all

Open-weights non-commercial base model with 12B parameters, serving as the foundation for the open-source fine-tuning ecosystem.

#19FLUX.1 [schnell]Black Forest Labs
Compare vs all

Apache 2.0 4-step distilled open weights model optimized for local inference and instant preview generation.

#20Ideogram 4.0 Open WeightsIdeogram
Compare vs all

Open-weight release of Ideogram 4.0, bringing world-class text rendering to self-hosted enterprise infrastructure.

#21Wan2.6 Text to ImageAlibaba
Compare vs all

Alibaba high-efficiency open video/image unified transformer backbone.