Interactive Utility
Qwen 3.8 Flash Next (180B) VRAM Calculator
Estimate the memory needed to run Qwen 3.8 Flash Next (180B) locally. Compare quantization levels and context lengths, then check the assumptions before choosing hardware.
How much VRAM does Qwen 3.8 Flash Next (180B) need? Qwen 3.8 Flash Next (180B) (180B parameters) needs about 118.6 GB of VRAM at Q4_K_M and about 369.5 GB at FP16, with an 8k-token context. Estimate from parameter count, bits per weight, KV cache and 1.5 GB of runtime overhead — not a benchmark.
Model
Custom: billion
Smaller modeled weight footprint; check model-specific quality.
Estimated memory footprint
- Weights (Q4_K_M)
- 109.1 GB
- KV cache (8k)
- 8.0 GB
- Runtime overhead
- 1.5 GB
- Estimated system RAM to load
- 111 GB
- Estimated model file size
- 114.6 GB
Compare with your available memory
Weights exceed available memory
The modeled weights alone exceed your usable memory. Try a smaller model, stronger quantization, or more memory.
Qwen 3.8 Flash Next (180B) memory requirements by quantisation
Estimated at an 8k-token context. Qwen 3.8 Flash Next (180B) is 180B parameters; the arithmetic and its assumptions are set out below.
| Quantisation | Weights | Total VRAM |
|---|---|---|
| FP16 (unquantised) | 360.0 GB | 369.5 GB |
| Q8_0 (8-bit) | 191.3 GB | 200.8 GB |
| Q6_K (6-bit) | 148.5 GB | 158.0 GB |
| Q5_K_M (5-bit) | 128.3 GB | 137.8 GB |
| Q4_K_M (4-bit) | 109.1 GB | 118.6 GB |
| Q3_K_M (3-bit) | 87.8 GB | 97.3 GB |
Arithmetic over stated assumptions, not a benchmark. Weights are parameter count times bits-per-weight; the KV cache term assumes grouped-query attention rather than this model's verified KV-head count; 1.5GB is reserved for runtime overhead. Confirm the model configuration and usable device memory before purchasing hardware.
Size a related model
Embed this Qwen 3.8 Flash Next (180B) Calculator
Embed a calculator prefilled with this model's parameter count in your blog, documentation, or internal wiki.
<iframe src="https://fathomlayer.com/embed/hardware-calculator?p=180" width="100%" height="600" frameborder="0" style="border-radius: 12px; border: 1px solid rgba(255,255,255,0.1);"></iframe>
How this math works
Local inference depends on several constraints. Memory capacity affects whether a model loads; memory bandwidth, compute and runtime settings also affect generation speed.
- Model Weights: At FP16 (unquantized), every 1 Billion parameters requires ~2GB of VRAM. At Q4_K_M (4-bit quantization), that drops to about 0.6GB per 1B parameters.
- KV Cache (Context Window): As you feed text into the model, it stores attention states in memory. A 32k context window on a 70B model requires several extra Gigabytes of RAM independent of the model weights.
- Overhead: This estimate reserves 1.5 GB for the runtime. Actual allocation depends on the device, framework and workload, so compare against usable rather than advertised memory.
