Model launches, price changes, and comparison analysis, with links to source-backed catalog facts and the models discussed.
Continue exploring independent model comparisons, hardware fit calculators, and live market movements.
Compare any two models on our category ratings, speed classes and token pricing.
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live audit trail of benchmark updates, new model releases, and API price cuts.
Answer 5 quick questions to compute deterministic model recommendations for your use case.
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.

Between 21 and 22 September, six models with a published Intelligence Index score reached OpenRouter. Claude Opus 5.5 takes the top of the index at 58. Xiaomi's MiMo-V2.6-Pro matches Grok 4.7 at a fifth of the output price, and Cohere's Command A+ returns a first token in 0.3 seconds.

OpenAI filled out the GPT-6 line on 22 September. Sol keeps GPT-5.6 Sol's $2/$10 price, scores one point higher and streams 37% faster. Luna halves input cost and cuts output cost by 58%, for one point less on the Intelligence Index.

xAI's Grok 4.7 lists at $1.60 input and $4.80 output per million tokens, 20% below Grok 4.6. It streams 27% faster, and it edges up from 44 to 46 on the Artificial Analysis Intelligence Index. Context and output ceilings are unchanged at 500K and 450K.

Anthropic's Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, five points clear of anything else in our catalog, and lists at $4 input and $20 output per million tokens. That is 20% below Opus 5 and 60% below Fable 5.1 and GPT-6 Astra.

Seventeen models from thirteen vendors reached the market in September, and the top of the Intelligence Index moved from 51 to 53. No vendor publishes a roadmap, so here is what the release record actually supports.

Claude Opus 4.8 and GLM-5.3 Flash both score 42 on the Artificial Analysis Intelligence Index. They were released 91 days apart. One bills $25.00 per million output tokens and the other bills $0.50.

Google's Gemini 3.1 Flash Image generates at up to 4096x4096 and bills $67 per 1,000 images at a 9.1-second median. It ranks seventh on the image arena — below three OpenAI models costing three times more.

Alibaba's Qwen-Audio-3.0-TTS Plus scores 1260 on the Artificial Analysis speech arena against 1168 for ElevenLabs' Eleven v3, and bills $27.60 per million characters against $100.00. It also speaks 16 languages.

OpenAI holds the top two places on the image arena. It also charges twenty-one times what the eighth-placed model costs. Here is the full quality-versus-price ladder, with the numbers.

Alibaba's Wan 3.0 sits 141 Elo above Google's Veo 3.1 on the Artificial Analysis video arena and bills $12.00 per minute against Veo's $24.00. The catch is latency, and it is a big one.