Skip to main content

Reported-benchmark ratings

Ratings labelled “reported” use an explicitly published admin evidence release. They are not community Elo or controlled head-to-head results. Models may use different benchmarks and test settings; numerical ordering is indicative, not proof of superiority. The effort-matched methodology below remains a separate legacy basis.

  1. Each benchmark has documented fixed minimum/maximum bounds and direction. Results are linearly normalized to 0–100; lower-is-better scales are inverted. Out-of-range results are rejected.
  2. Each category contains weighted benchmark families. Only one accepted result contributes per family: our own tests first, then vendor evidence, then third-party evidence; followed by the authored benchmark preference order, newest report date and stable ID. The highest achieved score is never the selection rule.
  3. Category score is the weighted mean of available family results. Displayed rating is round(1000 + 10 × score). Missing families do not count as zero; coverage is present family weight divided by applicable family weight. No evidence means unrated.
  4. Weights, reasons, sources, actual configurations, raw results, normalization and contributions are disclosed on the model page and comparison. Differing input sets remain visible; normalization does not establish scientific equivalence.
  5. Overall uses the declared category-strength weights and minimum category coverage; its contributing categories are disclosed. There is no combined cross-task audio Overall. Votes, price and speed do not enter this reported formula.
  6. Admins preview and explicitly publish immutable releases. Evidence corrections use new observation IDs; method changes are not claims of model improvement. Legacy API Elo fields retain their old meaning, with reported ratings provided separately.

What this methodology covers

  1. Bounded catalog and independent publication, tracking and SEO gates.
  2. Commercial facts vs effort-aware speed facts vs reference observations vs CompareLLM opinion.
  3. Category and Overall rating, effort-separated lists, recipes, coverage, votes after the threshold.
  4. Store exact / show class; exact speed lives in disclosure, methodology and changelog detail.
  5. Evidence tiers; `own_run` requires a package/checksum/protocol.
  6. Admin approval only. No automated catalog or benchmark ingest. Sync proposes; it never applies.
  7. Corrections append. Projections rebuild. Caches invalidate after an accepted write.
  8. Limits: pool-relative unitScores, coverage, source disagreement, hardware estimates, FX display.
  9. Ratings are our researched opinion from admissible evidence. Coverage and any profile are disclosed.
  10. Recipe weights change only after research, with a cited note and a new `recipeVersion`.
  11. Value, price, speed and context are fact boards with percentile — never a blended value score.

Research notes for the current recipe bind to recipe-v1. Profile notes will appear here when a model rating profile is published.

Evaluation Standards & Transparency

CompareLLM Methodology

CompareLLM publishes a bounded catalog and our own category Elo. We are not an automated aggregator. Admins file identity, facts and evidence, and every commercial fact and observation is stored as a timestamped snapshot carrying its source. Ratings compute at read time, and pages never fabricate or extrapolate a value we do not hold — missing stays missing. AI may draft news after approval; it never writes catalog, ratings or benchmarks.

Estimated ratings and improving coverage

When an applicable category has no accepted benchmark result, an admin may publish a 0–100 sentiment-based quality estimate with sources, rationale, confidence and a review date. The rating scale is 1000 + 10 × quality score. These are editorial judgments, not measured accuracy or win probabilities.

An estimate participates in category and applicable Overall rankings with an Estimated label and zero benchmark coverage. Accepted benchmark results take precedence after explicit publication. Prior estimates remain in release history. Confidence is an editorial assessment, separate from benchmark coverage. Unsupported tasks remain Not applicable, and audio tasks retain separate ratings rather than an invented cross-task Overall.

The two layers: measured and stated

Almost everything on this site is the first layer: a number someone else measured, recorded with its source and the date it was observed. Rankings, comparisons and recommendations are computed from those numbers. No human hand-places a model in them, and no provider can pay to move one.

The second layer is our opinion. Benchmarks do not measure everything that decides a real choice — how a model holds up over a long agent run, how its tooling behaves under load, whether a licence is workable. Where we have a view the numbers cannot show, we publish it as “CompareLLM’s Take”: attributed to us, dated, with the reason written out.

What an editorial pick never does

  • Change a metric, a snapshot, or the source and date attached to it.
  • Reorder a benchmark-sorted table or leaderboard.
  • Appear in the structured data that states our measured findings.
  • Appear without being labelled as an opinion.

How an editorial pick is kept honest

  • It records who made it and when, and states why in full.
  • It is written against a snapshot of the numbers at that moment.
  • It is re-checked against current data, and flags itself when the gap it relied on narrows.
  • It expires unless a person re-confirms it, and disappears on its own if a newer model overtakes the call.
  • It is never sold. No provider can pay for one.

If you only want the measurements, they are unchanged and always shown beside any opinion we add — including on the page you were reading when you followed this link.

Approval-gated 4-step evidence pipeline

On-demand · admin initiated
1Evidence research

An admin files the observation

A person records the exact value with its benchmark and version, publisher, source link, evidence tier, effort and observed date. No scraper, feed or AI writes a catalog fact, a benchmark result or a rating. An OpenRouter sync may propose a price or default-effort speed; an admin approves or rejects every proposal.

2Alias Normalization

Deterministic Name Resolution

Maps provider-specific IDs (e.g. anthropic/claude-3.5-sonnet:beta) to canonical catalog models using verified alias dictionaries.

3Rating computed at read

Category rating, one effort at a time

Each considered input is scored against the models on that one list, then blended by the disclosed recipe into a category Elo of roughly 1000–2000. A missing input is skipped and the coverage is shown — never filled in with a middle value. A model with no admissible evidence stays unrated rather than being ranked last.

4Changelog & Revalidation

Audit trail & invalidation

An accepted write is appended to /changelog with its before and after, and every list, page, feed and social card that depended on it is rebuilt. Corrections append and supersede; we do not edit a published observation in place.

Primary Data Sources

Provider catalog cross-check

List prices ($/1M input and output tokens), maximum context window limits, and prompt format specs.

Pricing & Discovery

LiveBench (Monthly Releases)

Contamination-resistant evaluation over mathematical reasoning, data analysis, and coding tasks.

Reasoning & Math

Princeton SWE-bench Verified

Real-world software engineering resolution rate across verified GitHub issue test harnesses.

Agentic Coding

LMSYS Chatbot Arena

Crowdsourced pairwise Bradley-Terry Elo preference rankings across coding and general conversations.

Human Preference Elo

Benchmark Data Provenance & Trust Tiers

Every score and metric snapshot on CompareLLM carries an explicit provenance tier so users can immediately distinguish direct empirical evaluations from manufacturer estimates.

Tier 1: Direct Benchmark Harness

Direct Read

Independently executed runs from recognized third-party test suites (LMSYS Arena, LiveBench, Princeton SWE-bench, VBench, GenAI-Bench). Linked directly to evaluator datasets.

Tier 2: Live Marketplace Feeds

Market Telemetry

Real-time on-demand pricing, token generation speed, time-to-first-token latency, and context window capacities are researched from provider and benchmark evidence, then reviewed by an administrator before publication.

Tier 3: Manufacturer Spec Baseline†

Spec Baseline

Preliminary manufacturer self-reported specs or bootstrap baselines for unreleased/new models. Always flagged with a † glyph until replaced by an independent test run.

Tier 0: Human & Agent Audited✓

Verified Audit

Manually audited against published evaluation papers and vendor technical reports with recorded cross-check dates and source references.

Transparency: What We Do vs What We Do Not Do

What We Do

  • • Record timestamped, dated snapshots for every single metric.
  • • Link directly to evaluator sources and harness versions.
  • • Keep full public audit logs of every score shift in /changelog.
  • • Render mathematical non-dominated Pareto convex hulls.

What We Do Not Do

  • • We do not scrape closed third-party portals.
  • • Editorial estimates are explicitly labelled; missing commercial facts and benchmark measurements are never fabricated.
  • • We do not accept sponsored placements or artificial rank boosting.
  • • We do not promote unverified models into indexable sitemaps.

CompareLLM rating, and how it differs from preference Elo

Our rating is a category Elo: one number for one job — Coding, Reasoning, Writing, Agents, Chat and so on — at one reasoning effort. It runs roughly 1000–2000, where 1500 is about the middle of the models we rate on that list. It is computed at read time from the inputs the category recipe names, so it changes when the evidence changes, not when someone edits a number.

Efforts are separate lists. A high-effort result never appears on the default board, and a rating is never copied between them — the same model can sit in a different place on each.

Preference Elo is something else entirely: LMArena publishes it from crowd pairwise votes. We may consider it as one reference among others, but it is not our rating and it never decides a board here.

A CompareLLM rating is our own verdict for one job — Coding, Reasoning, Writing and so on — at one reasoning effort. It runs roughly 1000 to 2000, like a chess rating: it places a model against the others on that list rather than scoring it out of 100. It is computed from a disclosed recipe of published evidence, each input carrying its source and date, and every rating says how much of that recipe it could actually use. It is our opinion, not an exam mark and not a crowd vote.

Longer explainer: What is Elo on an AI leaderboard? · Live Elo ranking

How the site updates (plain language)

  1. An administrator files benchmark evidence on the Evidence desk with its publisher, exact variant and version, and evidence tier. Nothing writes a benchmark value unattended.
  2. Commercial facts and default-effort speed can be fetched from OpenRouter on an explicit refresh. They arrive as proposals; an administrator approves or rejects each one, and a marketplace feed never displaces a better-sourced value.
  3. Ratings recompute from current evidence. Nobody types a public Elo — the only editorial lever is a disclosed own-strategy leg with a written reason.
  4. Model identities are matched through governed aliases and redirects. A new identity can be published but not indexed while its evidence coverage matures — publication, tracking and indexing are three separate decisions.
  5. When an accepted write moves something material we record it in /changelog and rebuild every page, feed and card that depended on it.
  6. A public login can vote, comment and save models. Filing evidence, approving proposals and publishing news are admin actions.

How CompareLLM updates every day · How a new model joins

What these ratings cannot tell you

  • A rating is relative to its pool. Each input is scored against the other rated models on that same list, so an Elo describes a position among the models we track, not an absolute capability.
  • Coverage varies between models. Two models on one board may be rated from different evidence. Every rating shows how much of its recipe was actually present, and a comparison says so when the two differ materially.
  • Sources disagree. Where a stronger source already holds a value, a weaker one does not replace it — it stays as history rather than overwriting a better-evidenced number.
  • Hardware figures are estimates. The hardware calculator models a device; it is not a measurement and never becomes a catalog speed fact.
  • Currency is display only. Prices are stored in USD as published. Changing currency converts what you see and never rewrites the fact.

There is no blended “value score” here. Price, speed and context are ordered as fact boards — one exact stored number with a position in its pool — because mixing a price into a capability rating produces a figure that answers neither question. Where value matters, use the price-versus-Elo view or a named decision preset, both of which say what they weighed.

Frequently asked questions

Plain-English methodology and leaderboard answers

A CompareLLM rating is our own verdict for one job — Coding, Reasoning, Writing and so on — at one reasoning effort. It runs roughly 1000 to 2000 and is computed from the inputs that category's published recipe names, so it moves when the evidence moves. It is not LMArena's preference Elo, which is a separate third-party number we may consider as one reference among several.
Platform Discovery & Next Steps

Explore More AI Intelligence Tools

Continue exploring independent model comparisons, hardware fit calculators, and live market movements.

Back to Homepage Overview