What is Elo on an AI leaderboard?
Elo is a crowd vote: people pick which hidden answer they like more. A higher Elo means more people preferred that model — not that it passed a test.
In one minute
Imagine two AIs write an answer to the same question. A person sees both answers with the names hidden. They tap the one they like more. Nobody is grading math homework. They are just saying “this one felt better.”
After thousands of those votes, the AI that wins more often gets a higher Elo. A higher Elo means people preferred its answers more often. It does not mean the model is smarter, cheaper, faster, or better at coding.
Where the number comes from
Elo started as a chess rating: when two players compete, the winner takes points from the loser. On this site the “players” are language models. The votes come from public LMArena / Arena tables. We copy the dated snapshot. We do not run the voting booth, and we do not invent a missing cell.
The column is labeled Preference Elo so it is not confused with a school test or a 0–100 “intelligence index.”
What Elo is not
It is not a science exam. A friendly, long answer can beat a short, correct one if voters like the style.
It is not the coding test (SWE-bench: did an agent finish a real GitHub issue?). It is not LiveBench (objective tasks). It is not “did it tell the truth?” Coding Elo, when we have it, is a separate vote on coding prompts — do not mix it with general Elo.
How to use it here
Use Elo to shortlist a general chat assistant. Then open a vs page and check the coding test, speed, and price for the actual job. The highest Elo model is often the wrong pick for a cheap widget or a repo agent.
If two models are close (tens of points, not hundreds), treat the crown as noise. Ratings move as new votes arrive. Always read the as-of date.
Why other sites quote a different Elo
Different dumps, different dates, and different name matching. We refuse to scrape a third-party “intelligence index.” If our cell is blank, the public feed did not match a name yet.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
Related guides
Same topic, next level of detail.
SWE-bench vs preference Elo
Why the best coding model is not always the best chatbot, and how CompareLLM keeps both numbers visible.
What is LiveBench?
Contamination-resistant objective tasks. Only compare scores from the same LiveBench release.
How to read LLM benchmarks without getting fooled
A short checklist: same split, same date, named source, no composite index you cannot audit.
Explore More AI Intelligence Tools
Pairwise Model Comparisons
Compare any two models on our category ratings, speed classes and token pricing.
Curated 'Best-Of' Indexes
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
AI Hardware Calculator
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live Market Intelligence
Live audit trail of benchmark updates, new model releases, and API price cuts.
Interactive Model Finder
Answer 5 quick questions to compute deterministic model recommendations for your use case.
Methodology & Standards
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.
