Free tools

LLM price comparison

Enter your monthly requests, token sizes, and cache hit rate. See what the same workload costs on every model in the catalog, cheapest first.

01Ox Alpha$0.00/mo
02Qwen 2 Instruct (1.5B)Open weight$11.18/mo
03LFM2.5-8B-A1BOpen weight$32.52/mo
04Meta Llama 3.2 1B InstructOpen weight$33.54/mo
05Meta Llama 3.2 3B InstructOpen weight$33.54/mo
06Trinity MiniOpen weight$43.53/mo
07Gemma 3N E4B InstructOpen weight$44.04/mo
08OpenAI GPT-OSS 20BOpen weight$54.20/mo
09Arize AI Qwen 2 1.5B InstructOpen weight$55.90/mo
10Nvidia Nemotron Nano 9B V2Open weight$66.79/mo
12DeepSeek R1 Distill Qwen 1.5BOpen weight$101/mo
14DeepSeek V4 FlashOpen weight$106/mo
15Qwen3.5 9B FP8Open weight$109/mo
16Llama Guard 4 12BOpen weight$112/mo
17Meta Llama 3 8B InstructOpen weight$112/mo
19Meta Llama 3.1 8BOpen weight$112/mo
20Ministral 3 14B Instruct 2512Open weight$112/mo
21Mistral (7B) Instruct v0.1Open weight$112/mo
22Mistral (7B) Instruct v0.3Open weight$112/mo
23OpenAI GPT-OSS 120BOpen weight$163/mo
24Qwen2.5 7B Instruct TurboOpen weight$168/mo
26Qwen3-VL-8B-InstructOpen weight$188/mo
27Pearl-ai Gemma-4-31B-it-pearlOpen weight$258/mo
28Glm 4.5 Air Fp8Open weight$269/mo
29GLM 4.5 AirOpen weight$272/mo
30GPT-5.6 Luna$287/mo
31Gemma 4 31B-it FP8Open weight$320/mo
32Qwen3 Next 80B A3b InstructOpen weight$320/mo
33Qwen3 Next 80B A3b ThinkingOpen weight$320/mo
34MiniMax M2.7 FP4Open weight$332/mo
35MiniMax M3$332/mo
36Mixtral-8x7B Instruct v0.1Open weight$335/mo
37Nous Hermes 2 Mixtral 8X7B DpoOpen weight$335/mo
38Qwen3.7 Plus$347/mo
39Qwen3 Coder Next Fp8Open weight$402/mo
40Deepseek Coder 33B InstructOpen weight$447/mo
41Gemma-2 Instruct (27B)Open weight$447/mo
42Qwen 2.5 14B InstructOpen weight$447/mo
43Qwen 2.5 Coder 32B InstructOpen weight$447/mo
44Qwen3-VL-32B-InstructOpen weight$455/mo
48Qwen2 72B InstructOpen weight$503/mo
49GLM 4.7 FP8Open weight$523/mo
50Deepseek V3.1 NVFP4Open weight$528/mo
52GLM 4.6 Fp8Open weight$615/mo
53Qwen QwQ-32BOpen weight$671/mo
54Qwen2-VL (72B) InstructOpen weight$671/mo
55Qwen2.5 72B InstructOpen weight$671/mo
56Qwen2.5 72B Instruct TurboOpen weight$671/mo
57Kimi K2.5 Fp4Open weight$682/mo
58Cogito v2.1 671BOpen weight$699/mo
59Qwen3.6 Plus$717/mo
62DeepSeek R1 Distill Qwen 14BOpen weight$894/mo
63Qwen3.5 397B A17bOpen weight$930/mo
67Grok 4.3$936/mo
68GLM 5 Fp4Open weight$944/mo
69Kimi K2.7 CodeOpen weight$1,088/mo
70Inkling FP4Open weight$1,110/mo
71DeepSeek R1 Distill Llama 70BOpen weight$1,118/mo
73Qwen3.7 Max$1,136/mo
74Kimi K2.6 Fp4Open weight$1,268/mo
75Deepseek V4 Pro$1,283/mo
76GLM 5.1 FP4Open weight$1,336/mo
77GLM 5.2Open weight$1,336/mo
78Grok 4.5$1,842/mo
79Gemini 3.6 Flash$1,889/mo
80Grok 4.6$1,890/mo
81Meta Llama 3.1 405B InstructOpen weight$1,957/mo
82Qwen2.5-VL (72B) InstructOpen weight$2,149/mo
83Gemini 3.5 Flash$2,151/mo
84DeepSeek R1 0528 NVFP4Open weight$2,377/mo
85Gemini 3.1 Pro$2,868/mo
86Claude Sonnet 5$3,777/mo
87Kimi K3Open weight$3,777/mo
88Claude Opus 4.8$6,295/mo
89GPT-5.5$7,170/mo

Same workload, 717000x gap between Ox Alpha and GPT-5.5. Route each task to the model that fits it.

Prices checked 21 Jul 2026 against published provider rates.

How it works

1
Describe the workload
Monthly requests, average input and output tokens per request, and your prompt cache hit rate.
2
Read the sorted list
Every language model in the catalog (89 today) priced on the same workload at provider list prices, cheapest first, with open-weight models marked.
3
Split the traffic
The gap between the cheapest and the most expensive model is usually 50 to 100x. Routing each task to the right model captures most of it.

Common questions

Where do the prices come from?

Provider list prices per million tokens from the Allocate catalog, split by input and output, with cached input priced at each provider’s published cache rate (or roughly 10% of input where none is published).

Why do output tokens cost more?

Generation is the expensive direction: every output token requires a full forward pass, while input tokens are processed in parallel. Most providers price output 2 to 6 times above input.

What is a prompt cache hit?

When the start of your prompt (system instructions, tool definitions) repeats across requests, providers can reuse the computed state and charge a fraction of the normal input price for those tokens. Agents with long stable system prompts see high hit rates.

Do I have to pick one model?

No. Most production teams route: judgment-heavy traffic goes to a frontier model, high-volume routine traffic to an open-weight one. On Allocate each route names its model, so you can swap either without a deploy when prices or quality move.

Is this calculator free?

Yes, free and unlimited, no account. It runs in your browser.

More free tools

These are provider list prices. On Allocate every model here sits behind one key, with per-token metering, hard caps, and a bill you forecast before you spend.