LLM price comparison
Enter your monthly requests, token sizes, and cache hit rate. See what the same workload costs on every model in the catalog, cheapest first.
Same workload, 717000x gap between Ox Alpha and GPT-5.5. Route each task to the model that fits it.
Prices checked 21 Jul 2026 against published provider rates.
How it works
Common questions
Where do the prices come from?
Provider list prices per million tokens from the Allocate catalog, split by input and output, with cached input priced at each provider’s published cache rate (or roughly 10% of input where none is published).
Why do output tokens cost more?
Generation is the expensive direction: every output token requires a full forward pass, while input tokens are processed in parallel. Most providers price output 2 to 6 times above input.
What is a prompt cache hit?
When the start of your prompt (system instructions, tool definitions) repeats across requests, providers can reuse the computed state and charge a fraction of the normal input price for those tokens. Agents with long stable system prompts see high hit rates.
Do I have to pick one model?
No. Most production teams route: judgment-heavy traffic goes to a frontier model, high-volume routine traffic to an open-weight one. On Allocate each route names its model, so you can swap either without a deploy when prices or quality move.
Is this calculator free?
Yes, free and unlimited, no account. It runs in your browser.
More free tools
These are provider list prices. On Allocate every model here sits behind one key, with per-token metering, hard caps, and a bill you forecast before you spend.