Skip to main content
CompareLLM Technical Deep-Dive
Updated 2026-08-16

How to read LLM prices ($/1M tokens)

List price is not your invoice. Cache, retries, long context, and output tokens move the real bill.

1Section 1

Two numbers, not one

Input $/1M is what you pay to send the prompt (and usually the RAG dump). Output $/1M is what you pay for the reply. Coding agents spend more on output. RAG spends more on input.

CompareLLM pair pages sketch three list-price scenarios. They are not a finance model. Cache hits, batch APIs, and long-context surcharges are not in the cell.

2Section 2

Why a “cheap” model can cost more

If it retries three times or dumps a 100k context on every turn, the cheaper row loses. Compare the pair page, then measure on your own traffic.

Ready to evaluate your stack?

Calculate your optimal model weights with Stack Engine or compare top models head-to-head.

Explore More AI Intelligence Tools

Back to Homepage Overview