How CompareLLM compares AI models
What Elo, LiveBench, SWE-bench, TTFT, and list price mean on this site — and what we refuse to invent.
One row is a snapshot, not a personality
Every cell on CompareLLM is a dated snapshot: a number, a unit, a source name, and an observed-at time. If we do not have a public machine-readable feed for a metric, the cell is blank. We do not paint radar axes with guessed 0–100 scores.
CompareLLM is not an aggregator. We publish our own category Elo — a rating we compute from a disclosed recipe of admissible evidence — alongside exact commercial facts and speed classes. We do not run SWE-bench or LiveCodeBench ourselves; those are references our recipes may consider, named on every rating.
The five numbers people actually argue about
CompareLLM category Elo is our rating for one job at one effort, roughly 1000–2000, computed from the inputs the recipe names. LMArena's preference Elo is a different, third-party number: crowd pairwise taste, not our rating and not a science exam.
LiveBench is a contamination-resistant objective suite. Only compare scores from the same release.
SWE-bench is “did this harness resolve a real GitHub issue?” Agent and split change the number.
TTFT and tok/s are latency and stream rate. List $/1M is not your invoice after cache and retries.
How a new vs page appears
We do not hand-write Claude vs GPT pages. If two models are indexable and share at least three metrics, /compare/{a}-vs-{b} exists (canonical A–Z slug). Frontier pairs are prerendered; the rest are created on first request.
A new model enters as a draft, and an administrator files its identity and evidence explicitly. Nothing is discovered or written automatically — there is no automated ingest and no AI writes a catalog value, a price or a rating.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
Related guides
Same topic, next level of detail.
Explore More AI Intelligence Tools
Pairwise Model Comparisons
Compare any two models on our category ratings, speed classes and token pricing.
Curated 'Best-Of' Indexes
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
AI Hardware Calculator
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live Market Intelligence
Live audit trail of benchmark updates, new model releases, and API price cuts.
Interactive Model Finder
Answer 5 quick questions to compute deterministic model recommendations for your use case.
Methodology & Standards
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.
