Model catalog

Open-weight models with the longest context windows

The longest-context open-weight models on the catalog reach 1M tokens: Llama Guard 4 12B leads at $0.20 per million input tokens. A 1M-token window holds roughly 750,000 words, an entire policy library or codebase in one prompt, on weights you can fine-tune and own.

01
Meta · 1M context · $0.20 in / $0.20 out per M tokens · Llama community
02
Deepseek · 1M context · $0.14 in / $0.28 out per M tokens · Not listed
03
Meta · 1M context · $0.18 in / $0.59 out per M tokens · Llama community
04
Moonshot AI · 1M context · $3 in / $15 out per M tokens · Not listed
05
Thinking Machines · 512K context · $1 in / $4.05 out per M tokens · Apache 2.0
06
Mistralai · 256K context · $0.20 in / $0.20 out per M tokens · Apache 2.0
07
Qwen · 256K context · $0.17 in / $0.25 out per M tokens · Not listed
08
Qwen · 256K context · $0.18 in / $0.68 out per M tokens · Apache 2.0
09
pearl.ai · 256K context · $0.28 in / $0.86 out per M tokens · Not listed
10
Google · 256K context · $0.39 in / $0.97 out per M tokens · Apache 2.0

Provider list prices from the Allocate catalog, checked 2026-07-21.

What long context is worth

Long context replaces retrieval plumbing for bounded corpora: instead of chunking and fetching fragments, the model reads the whole source. The tradeoff is per-request cost, which grows with the tokens actually processed, and memory on the serving side, where the KV cache grows linearly with context.

Open weights change the economics of long-document fine-tuning too: training examples that are whole documents need a window that holds them, which this list ranks directly.

Common questions

What is the longest context window on an open model?

1M tokens, on Llama Guard 4 12B, at $0.20 per million input tokens at list.

Does long context cost more?

Per token, no: you pay the same list price per million tokens. Per request, yes, because you send more tokens. A full 1M-token prompt on a $0.18 model costs about $0.18 at list before caching.

Every model here sits behind one key on Allocate: route by name, meter per route, and swap the model in one click.