Model catalog

DeepSeek V4 Flash

Open weights

DeepSeek V4 Flash is a language model from Deepseek with a 1M-token context window. Provider list price is $0.14 per million input tokens and $0.28 per million output; on Allocate you pay $0.15 and $0.30. The weights are open, so you can fine-tune it and own the result.

Pricing

Provider listOn Allocate
Input, per M tokens$0.14$0.15
Output, per M tokens$0.28$0.30
Cached input, per M tokens$0.028$0.03

Prices checked 2026-07-21.

Price against its peers

DeepSeek V4 Flash$0.14$0.28
Qwen3.5 9B FP8$0.17$0.25
InputOutput

Provider list prices per M tokens, DeepSeek V4 Flash against its nearest language peers by price.

What a real workload costs

Take 1,000,000 requests a month at 1,200 input and 350 output tokens each: 1,200M input and 350M output tokens. At list prices that is 1,200 × $0.14 + 350 × $0.28 = $266 a month. Billed on Allocate it is $284.62.

DeepSeek V4 Flash is an open-weights model; the catalog does not list its license, so check the lab’s model card before commercial fine-tuning.

Example usage

Point a route at deepseek/deepseek-v4-flash and the endpoint never changes; swap the model behind it whenever you want.

api.allocate.network
curl https://api.allocate.network/v1/chat/completions \
  -H "Authorization: Bearer $ALLOCATE_KEY" \
  -d '{
    "model": "deepseek/deepseek-v4-flash",
    "messages": [{"role": "user",
      "content": "Summarise the attached contract."}]
  }'
200 · deepseek/deepseek-v4-flash · inside your boundary

Common questions

How much does DeepSeek V4 Flash cost per million tokens?

Provider list price is $0.14 per million input tokens and $0.28 per million output tokens. On Allocate you pay $0.15 in and $0.30 out.

What context window does DeepSeek V4 Flash have?

1,048,576 tokens (1M). At roughly 0.75 words per token, that is about 786k words of English text per request.

What does cached input cost on DeepSeek V4 Flash?

$0.028 per million tokens at list ($0.03 billed). Repeated prompt prefixes, such as a stable system prompt or tool definitions, bill at this rate instead of the full input price.

Can I fine-tune DeepSeek V4 Flash?

Yes. DeepSeek V4 Flash is an open-weights model; check the lab’s model card for the exact license terms. Read the license terms before fine-tuning for commercial use. On Allocate the trained weights stay inside your boundary and belong to you.

How do I call DeepSeek V4 Flash on Allocate?

Send deepseek/deepseek-v4-flash in the model field of the OpenAI-compatible endpoint at api.allocate.network/v1, or point a route name (like prod/support-agent) at it so you can swap the model later without a deploy.

Compare against