How to read LLM prices ($/1M tokens)
List price is not your invoice. Cache, retries, long context, and output tokens move the real bill.
Two numbers, not one
Input $/1M is what you pay to send the prompt (and usually the RAG dump). Output $/1M is what you pay for the reply. Coding agents spend more on output. RAG spends more on input.
CompareLLM pair pages sketch three list-price scenarios. They are not a finance model. Cache hits, batch APIs, and long-context surcharges are not in the cell.
Why a “cheap” model can cost more
If it retries three times or dumps a 100k context on every turn, the cheaper row loses. Compare the pair page, then measure on your own traffic.
Ready to evaluate your stack?
Calculate your optimal model weights with Stack Engine or compare top models head-to-head.
Related guides
Same topic, next level of detail.
Explore More AI Intelligence Tools
Pairwise Model Comparisons
Compare any two models on our category ratings, speed classes and token pricing.
Curated 'Best-Of' Indexes
Targeted rankings for Best Coding LLMs, Best Cheap APIs, Shortest Wait, and top-tier models.
AI Hardware Calculator
Interactive VRAM calculator, quantization levels (FP16, Q8, Q4), KV cache context, and local hardware fit.
Live Market Intelligence
Live audit trail of benchmark updates, new model releases, and API price cuts.
Interactive Model Finder
Answer 5 quick questions to compute deterministic model recommendations for your use case.
Methodology & Standards
How a CompareLLM rating is computed, what evidence it admits, and how prices and speed classes are recorded.
