How CompareLLM updates every day
Plain-language walkthrough of how facts and evidence are filed, reviewed and published.
One job, once a day
An administrator records model facts and benchmark observations through the Catalog and Evidence desks. AI suggestions remain proposals and never write rankings without approval.
If nothing material moved, the site stays as it is. If an administrator approves a material price, score, speed, or new-model change, we write a changelog row and revalidate the affected pages.
Where the numbers come from
Commercial facts and default-effort speed can be fetched from OpenRouter on an explicit refresh, and arrive as proposals. An administrator approves or rejects each one; a marketplace feed never displaces a better-sourced value, and it never touches identity, release dates, aliases or benchmark results.
Benchmark results are filed separately on the Evidence desk with their publisher, exact variant and version, evidence tier, and — for a first-party run — its result package. An unmatched OpenRouter id becomes a Catalog candidate, never a new model, so one marketplace listing cannot create forty thin compare pages overnight.
What you still do by hand
Publishing /news posts. AI may draft one after an administrator approves that generation, but the draft is private and a separate publish action makes it live. Approving the generation is not publishing it.
Every rating, too, in the sense that matters: the recipe behind each category Elo is authored and versioned by a person, and changing an input or a weight requires a cited research note.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
Related guides
Same topic, next level of detail.
Explore More AI Intelligence Tools
Pairwise Model Comparisons
Compare any two models on our category ratings, speed classes and token pricing.
Curated 'Best-Of' Indexes
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
AI Hardware Calculator
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live Market Intelligence
Live audit trail of benchmark updates, new model releases, and API price cuts.
Interactive Model Finder
Answer 5 quick questions to compute deterministic model recommendations for your use case.
Methodology & Standards
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.
