Free tools

Qwen3.5 397B A17b GPU requirements

What it takes to run Qwen3.5 397B A17b locally: memory by quantization, the smallest GPU that fits, and the managed alternative.

Qwen3.5 397B A17b serves up to 262,144 tokens of context; the KV cache grows linearly toward that ceiling, so the slider below shows exactly what longer context costs in memory.

FP16877 GB needed at 8K contextMulti-GPU
8-bit439 GB needed at 8K contextMac M3 Ultra
4-bit241 GB needed at 8K contextMac M3 Ultra
242 GBmemory needed · Qwen3.5 397B A17b at 4-bit, 16K context

Smallest single device that fits: Mac M3 Ultra (512 GB unified)

RTX 4090 (24 GB)11x needed
RTX 5090 (32 GB)8x needed
L40S (48 GB)6x needed
A100 (80 GB)4x needed
H100 (80 GB)4x needed
H200 (141 GB)2x needed
B200 (192 GB)2x needed
Mac M4 Max (128 GB unified)2x needed
Mac M3 Ultra (512 GB unified)Fits

Or skip the hardware: run Qwen3.5 397B A17b on Allocate, token-metered.

How it works

1
Check the table
Qwen3.5 397B A17b at 8K context across FP16, 8-bit, and 4-bit, with the smallest single device that fits each.
2
Tune the calculator
Longer context grows the KV cache. The slider shows exactly how much memory that adds.
3
Decide how to run it
MoE weights must fit in memory in full, so serving this locally means a multi-GPU fleet, not one card.

Common questions

How much VRAM does Qwen3.5 397B A17b need?

At 8K context: roughly 877 GB at FP16, 439 GB at 8-bit, and 241 GB at 4-bit, including KV cache and runtime overhead. Longer context adds memory linearly.

Can a single GPU run Qwen3.5 397B A17b?

At 4-bit, yes: a Mac M3 Ultra (512 GB unified) or larger handles it at 8K context. At FP16 you need multiple devices.

Why does Qwen3.5 397B A17b need so much memory as a MoE model?

Only some experts activate per token, which sets speed, but all 397 billion parameters must sit in memory. MoE trades memory for throughput.

What license is Qwen3.5 397B A17b under?

Apache 2.0. Permissive: run it, fine-tune it, and own the result.

Is there an alternative to buying GPUs for Qwen3.5 397B A17b?

Yes: managed inference. Allocate serves Qwen3.5 397B A17b token-metered from $0.6 per million input tokens at provider list price. Idle time costs nothing.

More free tools

Allocate serves open-weight models like Qwen3.5 397B A17b token-metered inside your own boundary. No hardware to buy, and if you fine-tune on your data, the weights are yours.