CompareLLM
CompareLLM.ai
Live
Leaderboard
Models
Compare
Best of & Stacks
Research & News
…
CompareLLM
CompareLLM.ai
Precision Benchmarks

Programmatic, dated AI model benchmarks, head-to-head comparisons, and Stack Engine presets.

Daily ingest · 06:00 UTC

Analytics & Benchmarks

  • AI Model Leaderboard
  • Head-to-Head Compare Hub
  • Models Directory
  • Stack Engine Presets
  • Frontier Models
  • Open Weights Catalog

Guides & Intent Lists

  • Best LLM Lists (2026)
  • Best Coding LLM
  • Best Cheap LLM
  • Fastest Low-Latency LLM
  • Claude vs GPT Benchmark
  • What is Elo?
  • Methodology Guides
  • News & Dispatches

Transparency & API

  • Evaluation Methodology
  • Benchmark Changelog
  • Public JSON API
  • llms.txt Specification
  • Privacy Policy
  • Sign In / Account

© 2026 CompareLLM. Public benchmark data aggregated from Arena Elo, LiveBench, SWE-bench & OpenRouter.

Every score has a dated snapshot.

Theme:
Currency:
  1. Home
  2. News
  3. Mistral Large 3: the European flagship with function calling
Mistral Large 3: the European flagship with function calling
launchVerified Dispatch
CompareLLM Intelligence Desk·Jan 18, 2026

Mistral Large 3: the European flagship with function calling

Jan 2026. Not the Elo crown. Useful when residency and tools matter more than Arena.

Search intent: mistral large 3 benchmark

Evaluated Models (3):
Mistral Large 3Mistral
Codestral 25.01Mistral
Claude Opus 5Anthropic
Share Analysis:
WhatsAppTelegramXLinkedInReddit
Verified Benchmark Scorecard3 Models Evaluated

Mistral Large 3 & Frontier Peer Benchmark Comparison

Live benchmark scores, throughput speeds, and token pricing with dynamic peer comparison.

LMSYS Arena Elo1,456Mistral Large 3
SWE-bench Verified55.4%Coding resolve %
Throughput Speed112 tok/sGeneration rate
Output Price / 1M$6.00List API rate
Interactive Model Comparison:3 models selected
Mistral Large 3ArticleCodestral 25.01ArticleClaude Opus 5Article
LMSYS Arena EloOverall human preference
quality
Mistral Large 3
1,456
Codestral 25.01
1,320
Claude Opus 5Top
1,624
Coding EloProgramming preference
quality
Mistral Large 3
—
Codestral 25.01
1,402
Claude Opus 5Top
1,618
SWE-bench VerifiedGitHub issue resolve %
quality
Mistral Large 3
55.4%
Codestral 25.01
51.8%
Claude Opus 5Top
79.2%
GPQA DiamondPhD-level science reasoning %
quality
Mistral Large 3
74.1%
Codestral 25.01
58.4%
Claude Opus 5Top
87.1%
LiveBenchContamination-free reasoning
quality
Mistral Large 3
61.8%
Codestral 25.01
54%
Claude Opus 5Top
74.8%
Throughput (Speed)Output tokens / second
speed
Mistral Large 3
112 tok/s
Codestral 25.01Top
148 tok/s
Claude Opus 5
72 tok/s
Time-to-First-TokenInitial latency (ms)
speed
Mistral Large 3
190 ms
Codestral 25.01Top
150 ms
Claude Opus 5
310 ms
Output Token Price$ per 1M reply tokens
price
Mistral Large 3
$6/1M
Codestral 25.01Top
$0.9/1M
Claude Opus 5
$50/1M
Input Token Price$ per 1M prompt tokens
price
Mistral Large 3
$2/1M
Codestral 25.01Top
$0.3/1M
Claude Opus 5
$10/1M
Context WindowMax token capacity
capacity
Mistral Large 3
128k
Codestral 25.01
256k
Claude Opus 5Top
1M
Benchmark / Metric
Mistral Large 3Mistral · Article Model
Codestral 25.01Mistral · Article Model
Claude Opus 5Anthropic · Article Model
LMSYS Arena EloOverall human preference
1,456
1,320
1,624Top
Coding EloProgramming preference
—
1,402
1,618Top
SWE-bench VerifiedGitHub issue resolve %
55.4%
51.8%
79.2%Top
GPQA DiamondPhD-level science reasoning %
74.1%
58.4%
87.1%Top
LiveBenchContamination-free reasoning
61.8%
54%
74.8%Top
Throughput (Speed)Output tokens / second
112 tok/s
148 tok/sTop
72 tok/s
Time-to-First-TokenInitial latency (ms)
190 ms
150 msTop
310 ms
Output Token Price$ per 1M reply tokens
$6/1M
$0.9/1MTop
$50/1M
Input Token Price$ per 1M prompt tokens
$2/1M
$0.3/1MTop
$10/1M
Context WindowMax token capacity
128k
256k
1MTop
Executive Key Takeaway
Jan 2026. Not the Elo crown. Useful when residency and tools matter more than Arena.

Mistral Large 3 is the European flagship in this catalog, with strong function calling and a mid list price. It will not win a global Elo sort against Opus 5 or Sol. That is fine. People search it for residency, EU procurement, and tool use. We keep a full snapshot so those queries land on a sourced page, not a brochure.

Function calling is not a scored axis

We do not invent a tools-bench. If that is the job, use your own traces. Elo is a weak prior for tool loops.

Empirical Evaluation & Architectural Analysis

Empirical evaluations recorded across the CompareLLM Leaderboard highlight how parameter scaling and inference optimization impact production throughput and cost-per-token economics. Reference the interactive scorecard above for verified snapshots or inspect the dedicated mistral-large-3 dossier.

Codestral is the code sibling

Codestral 25.01 is the specialist. Large 3 is the generalist. Don’t buy Large to save money on CI — look at Codestral and Flash-class rows.

Strategic Deployment Recommendation

  • Production Workloads: For enterprise workloads requiring strict reliability, cross-reference Top Coding LLMs and Cheapest High-Quality LLMs.
  • Direct Pair Comparison: Check model specifications at mistral-large-3 specs.
  • Architectural Stacks: Recommended architectural configurations can be evaluated on the CompareLLM Stack Engine.

Frequently Asked Questions

Is this an CompareLLM lab test score? No. Linked model dossiers and compare showdowns reflect dated snapshots gathered from named public evaluation benchmarks. Seed catalog entries remain transparently timestamped until daily ingest updates them.

Where can I see live comparisons for this model? Explore verified model specs at mistral-large-3.

What primary search query does this briefing answer? mistral large 3 benchmark.

Share Analysis:
WhatsAppTelegramXLinkedInReddit

Head-to-Head Showdowns for Mentioned Models

Featured Article ShowdownMistral Large 3 vs Codestral 25.01
Compare Benchmarks
Benchmark MatchupMistral Large 3 vs Claude Opus 4.5
Benchmark MatchupMistral Large 3 vs Claude Opus 4.6
Benchmark MatchupCodestral 25.01 vs Claude Opus 4.5
Benchmark MatchupCodestral 25.01 vs Claude Opus 4.6
Benchmark MatchupClaude Opus 5 vs GPT-5
Benchmark MatchupClaude Opus 5 vs GPT-5 mini

Related news

  • compare · Aug 10, 2026

    Sonnet vs Opus in 2026: same family, different invoice

  • compare · Jul 24, 2026

    Claude vs Gemini: coding, Elo, and output price — not a brand mashup

  • compare · Jul 24, 2026

    Claude vs GPT benchmark: this hub always remaps to the current Elo leaders

  • launch · Jul 24, 2026

    Claude Opus 5 is Anthropic’s default flagship — how to read it on CompareLLM

Back to All DispatchesExplore Comparison Matrix