Kimi K2.7 Code vs Inkling FP4
On provider list prices, Kimi K2.7 Code costs $0.95 per million input tokens against $1 for Inkling FP4: 1.1x apart. Output is $4 against $4.05.
Specifications and provider list prices from the Allocate catalog, checked 2026-07-21.
What the numbers say
Take 1,000,000 requests a month at 1,200 input and 350 output tokens each. That workload costs $2,540 a month on Kimi K2.7 Code and $2,618 on Inkling FP4 at list: a gap of $77.50.
Inkling FP4 reads 512K tokens per request against 256K for Kimi K2.7 Code, 2.0x the window. That decides which one can take whole documents without splitting them.
Choose Kimi K2.7 Code for
- The lower list price ($0.95 in / $4 out per M tokens)
Choose Inkling FP4 for
- The longer context window (512K vs 256K tokens)
- Fine-tuning under a permissive license (Apache 2.0)
Common questions
Which is cheaper, Kimi K2.7 Code or Inkling FP4?
Kimi K2.7 Code, on this workload shape. At list prices it is $0.95/$4 per million tokens in and out against $1/$4.05 for Inkling FP4. Billed on Allocate: $1.02/$4.28 against $1.07/$4.33.
Which has the bigger context window?
Inkling FP4: 524,288 tokens (512K) against 262,144 (256K) for Kimi K2.7 Code.
Can I fine-tune Kimi K2.7 Code or Inkling FP4?
Both publish open weights (Kimi K2.7 Code: Not listed; Inkling FP4: Apache 2.0), so both can be fine-tuned. On Allocate the trained weights stay inside your boundary and belong to you.
Related comparisons
Run the numbers on your workload
Or don’t choose. On Allocate a route name is the contract: point yours at one model today, swap to the other tomorrow, and compare them on your live traffic with per-token metering.