Qwen QwQ 32B is a Alibaba open-weight frontier model. Latest preference Elo in this catalog is 1,495 (seed-bootstrap (Aug 1, 2026)). SWE-bench sits at 67.5% (seed-bootstrap (Aug 1, 2026)). List output price is $1.2/1M. Context window is 131k tokens. Alibaba specialized open reasoning model competing with frontier closed reasoning models. Numbers below are dated snapshots, not a guarantee on your traffic mix.
Compare Qwen QwQ 32B against
Currently comparing Qwen QwQ 32B. Alternative family tiers available.
Simulate accuracy gains vs added response time & token cost
Reasoning models “think before answering” by generating internal reasoning tokens. Higher effort improves math, coding, and logic accuracy, but increases response delay and token costs.
Plotted against all active catalog models (50th percentile = catalog median).
Ranks in the top tier (≥75th percentile) for Math.
Dated snapshot metrics aggregated from official evaluators and API providers.
| Benchmark Metric | Reported Score | Observed Source |
|---|---|---|
| Preference Elo | 1,495 | seed-bootstrap · Aug 1, 2026 |
| Coding Elo | 1,520 | seed-bootstrap · Aug 1, 2026 |
| LiveBench | 68% | seed-bootstrap · Aug 1, 2026 |
| SWE-bench | 67.5% | seed-bootstrap · Aug 1, 2026 |
| GPQA Diamond | 83.2% | seed-bootstrap · Aug 1, 2026 |
| Time to first token | 320 ms | seed-bootstrap · Aug 1, 2026 |
| Output speed | 68 tok/s | seed-bootstrap · Aug 1, 2026 |
| Input price | $0.3/1M | seed-bootstrap · Aug 1, 2026 |
| Output price | $1.2/1M | seed-bootstrap · Aug 1, 2026 |
| Context window | 131k | seed-bootstrap · Aug 1, 2026 |
Search or pick any catalog row. Suggested matchups first, then the full list.
Anthropic current default flagship. Official API $5/$25 per 1M tokens and a 1M context window (Anthropic, Jul 24 2026).
Anthropic top-tier X-High reasoning model engineered for heavy multi-step tasks. Official API $10/$50 per 1M tokens (Claude Platform pricing, Aug 2026).
OpenAI 5.6 flagship tier. Official API $5/$30 per 1M tokens (OpenAI pricing, Jul 30 2026 update left Sol unchanged).
Prior Opus generation still billed at $5/$25. Kept as a compare baseline against Opus 5.
xAI Aug 12 2026 post-training refresh of Grok 4.5. Same $2/$6 API price, 500k context, stronger agentic traces.
Follow-on Opus release. Slightly behind 4.5 on official SWE-bench bash-only in the last published sweep; stronger multi-step autonomous agent reasoning.
Current Google Pro-class multimodal model. Long context, strong coding, billed like the 3.x Pro tier.
Anthropic frontier coding and computer-use model. SWE-bench leader on the official mini-SWE-agent harness in the Feb 2026 refresh.
OpenAI cost-efficient reasoning model with adjustable thinking effort and top coding benchmark scores.
OpenAI flagship reasoning model for 2025–26. Strong general preference Elo and multimodal coverage.
Z.ai Aug 14 2026 post-train of the GLM-5.2 744B base. Coding-plan live; open weights promised after a two-week safety review (z.ai/blog/glm-5.3).
Anthropic workhorse. Official $2/$10 per 1M tokens made permanent on Aug 10 2026.
Google frontier multimodal model with a multi-million-token context window.
OpenAI 5.6 mid tier. Official API $2/$12 per 1M after the Jul 30 2026 price cut.
OpenAI largest dense non-reasoning model with expansive world knowledge and reduced hallucinations.
Open-weight-adjacent DeepSeek flagship. High reasoning density per dollar.
Google Aug 13 2026 workhorse. Official intro price $0.75/$3.75 per 1M through Dec 31 2026; 1,048,576-token context (Google blog).
Moonshot AI frontier flagship reasoning and agent swarm model with 256k context and top-tier SWE-bench coding capability.
Workhorse Anthropic model: most of Opus coding quality at a mid-tier price.
Alibaba flagship. Strong math and multilingual code.
Moonshot flagship reasoning model with high-fidelity coding, deep thinking, and tool use across a 200k context.
OpenAI pioneer reasoning foundation model designed for complex science, math, and multi-step reasoning.
Google Jul 21 workhorse. Now shares the 3.7 Flash introductory $0.75/$3.75 rate through Dec 31 2026.
Fast Google frontier-adjacent model. Near-top official SWE-bench at a fraction of Opus price.
Zhipu flagship. Strong Chinese/English coding and agents.
Google experimental reasoning model that visualizes thoughts in real-time.
DeepSeek Jul 31 2026 price-performance SKU. Public listings put it near $0.14/$0.28 per 1M tokens.
First Claude 4 Opus generation. Baseline for opus 4 vs opus 5.
MiniMax coding model. Tied near the top of official SWE-bench bash-only in Feb 2026.
01.AI ultra-fast reasoning model delivering top LiveBench efficiency.
Zhipu high-throughput multimodal vision model with sub-second latency and competitive coding.
ByteDance Seed 2.1 Turbo, listed on public model timelines as an Aug 10 2026 API drop.
ByteDance flagship enterprise model with ultra-low token cost and 128k context.
Meta natively multimodal open-weight flagship.
OpenAI 5.6 fast/cheap tier. Official API $0.20/$1.20 per 1M after the Jul 30 80% Luna cut.
Moonshot foundation model with strong native bilingual reasoning and 65.8% SWE-bench Verified coding baseline.
Mistral European flagship with strong function calling.
First Sonnet 4 generation. Bridge between 3.5/3.7 and Sonnet 5.
Baidu current multimodal enterprise foundation model with broad Chinese knowledge.
Alibaba balanced flagship API model with high-throughput general reasoning.
OpenAI high-speed, cost-effective reasoning model optimized for STEM, math, and code generation.
Open-weight Qwen3 mixture-of-experts.
Anthropic cheap/fast Claude SKU. Official Claude Platform list price $1/$5 per 1M tokens (Anthropic Haiku page).
01.AI full-scale dense model for complex instruction following.
Alibaba dedicated open-weight code generation model with near-frontier SWE-bench Verified coding capability.
Moonshot ultra-long context model supporting up to 2 million tokens per request.
Meta open-weight Llama 4 long-context sibling of Maverick. Common public API lists sit near $0.08–$0.30 / $0.30–$0.70 per 1M; we store a conservative hosted list until OpenRouter overwrites.
Cost-efficient GPT-5 distill for high-volume agents.
Previous xAI generation before Grok 4. Kept for grok 3 vs grok 4.
Alibaba open-weight 27B drop dated Aug 14 2026. Dense enough to self-host; not a frontier MoE.
Meta flagship open-weight 405B dense foundation model with 128k context window.
DeepSeek open-weight Mixture-of-Experts coding model supporting 338 programming languages and 128k context.
Hybrid-reasoning Sonnet from 2025. Kept for historical compare pages.
ByteDance high-speed lightweight model priced at sub-cent levels.
Open-weights reasoning model trained with large-scale RL.
Previous Google long-context flagship.
Cohere enterprise RAG and tool-use model.
Previous OpenAI flagship. Still a common compare baseline on legacy pages.
Mistral code-specialist model.
Prior DeepSeek flagship. Baseline for v3 vs v4.
Previous Google speed workhorse.
Previous Meta 70B open-weight workhorse.
2024 workhorse. Still a high-intent compare against GPT-4o.
Legacy small OpenAI model. Useful as a cheap baseline.
First million-token Gemini Pro. Baseline for 1.5 vs 2.5 vs 3.x Pro.
GPT-4 Turbo 128k. Historical flagship for gpt-4 turbo vs gpt-4o / gpt-5.
Original Claude 3 flagship. Kept so opus 3 vs later Opus and vs GPT-4o still resolve.
Llama 3.1 70B instruct. Predecessor to 3.3 70B and Llama 4.
Previous cheap Claude. Haiku 4.5 is the current $1/$5 SKU.
OpenAI flagship diffusion-transformer image synthesis model with supreme prompt adherence, realistic textures, and complex text composition.
Fast distilled tier of GPT Image 2 designed for high-throughput interactive creative workflows at 70% lower price.
Previous OpenAI image generation flagship. High fidelity with proven enterprise reliability.
Legacy OpenAI image model benchmark baseline. Retained for historical comparisons.
State-of-the-art cinematic image synthesis foundation model known for photorealistic lighting and aesthetic composition.
Google DeepMind frontier multimodal image generation flagship with Deep Research and complex multi-object spatial reasoning.
Google high-speed multimodal generative model delivering top Elo performance at half the latency and cost of Pro.
Sub-1.5 second ultra-low latency tier for real-time applications and mobile game asset pipelines.
Black Forest Labs maximum-capacity flow-matching model with photoreal anatomy and leading typography fidelity.
Flexible step-distilled variant of FLUX.2 offering 99% of Max quality with 45% faster generation speeds.
First-generation professional API model from Black Forest Labs, renowned for 6x faster generation than original FLUX.1 Pro.
Open-weights non-commercial base model with 12B parameters, serving as the foundation for the open-source fine-tuning ecosystem.
Apache 2.0 4-step distilled open weights model optimized for local inference and instant preview generation.
Industry-leading typography, graphic design, and in-image text layout model with precise kerning and complex banner generation.
Open-weight release of Ideogram 4.0, bringing world-class text rendering to self-hosted enterprise infrastructure.
Specialized generative model for vector art, icon sets, illustrations, and commercial design assets with native brand color matching.
High-speed utility tier of Recraft V4.1 tailored for rapid asset production at $0.035/image.
High-resolution real-time generation model tuned for artistic composition and dynamic concept art.
Sub-2 second interactive generation engine delivering ultra-affordable $15/1k image generation.
Alibaba multimodal image synthesis foundation model with strong bilingual Chinese/English typography and cultural asset fidelity.
Alibaba high-efficiency open video/image unified transformer backbone.
ByteDance flagship image generation model with high photorealism and fine facial structure rendering.
High-value ByteDance image foundation model at $30/1k images.
Midjourney v7 premier creative image generation model with unparalleled stylistic nuances, coherent hands/anatomy, and cinematic grading.
Previous standard in aesthetic digital art generation, retained as a historical compare baseline.
Microsoft AI proprietary foundation image model for Copilot Studio and enterprise creative workflows.
Cost-optimized Microsoft AI image model delivering sub-2 second responses at $20/1k images.
HiDream high-definition visual generator optimized for photorealism and accurate complex lighting.
Luma Labs universal 3D-aware image synthesis engine with spatial geometry consistency.
Plain-English methodology and leaderboard answers
Elo is a crowd vote on which hidden answer people liked more — not a school test. What is Elo?