Skip to content

Interactive Utility

Qwen 3.8 Flash Next (180B) VRAM Calculator

Estimate the memory needed to run Qwen 3.8 Flash Next (180B) locally. Compare quantization levels and context lengths, then check the assumptions before choosing hardware.

How much VRAM does Qwen 3.8 Flash Next (180B) need? Qwen 3.8 Flash Next (180B) (180B parameters) needs about 118.6 GB of VRAM at Q4_K_M and about 369.5 GB at FP16, with an 8k-token context. Estimate from parameter count, bits per weight, KV cache and 1.5 GB of runtime overhead — not a benchmark.

Model

Custom: billion

Smaller modeled weight footprint; check model-specific quality.

4k32k128k

Estimated memory footprint

118.6GB
Weights (Q4_K_M)
109.1 GB
KV cache (8k)
8.0 GB
Runtime overhead
1.5 GB
Estimated system RAM to load
111 GB
Estimated model file size
114.6 GB

Compare with your available memory

Weights exceed available memory

The modeled weights alone exceed your usable memory. Try a smaller model, stronger quantization, or more memory.

Qwen 3.8 Flash Next (180B) memory requirements by quantisation

Estimated at an 8k-token context. Qwen 3.8 Flash Next (180B) is 180B parameters; the arithmetic and its assumptions are set out below.

QuantisationWeightsTotal VRAM
FP16 (unquantised)360.0 GB369.5 GB
Q8_0 (8-bit)191.3 GB200.8 GB
Q6_K (6-bit)148.5 GB158.0 GB
Q5_K_M (5-bit)128.3 GB137.8 GB
Q4_K_M (4-bit)109.1 GB118.6 GB
Q3_K_M (3-bit)87.8 GB97.3 GB

Arithmetic over stated assumptions, not a benchmark. Weights are parameter count times bits-per-weight; the KV cache term assumes grouped-query attention rather than this model's verified KV-head count; 1.5GB is reserved for runtime overhead. Confirm the model configuration and usable device memory before purchasing hardware.

Size a related model

Embed this Qwen 3.8 Flash Next (180B) Calculator

Embed a calculator prefilled with this model's parameter count in your blog, documentation, or internal wiki.

<iframe src="https://fathomlayer.com/embed/hardware-calculator?p=180" width="100%" height="600" frameborder="0" style="border-radius: 12px; border: 1px solid rgba(255,255,255,0.1);"></iframe>

How this math works

Local inference depends on several constraints. Memory capacity affects whether a model loads; memory bandwidth, compute and runtime settings also affect generation speed.

  • Model Weights: At FP16 (unquantized), every 1 Billion parameters requires ~2GB of VRAM. At Q4_K_M (4-bit quantization), that drops to about 0.6GB per 1B parameters.
  • KV Cache (Context Window): As you feed text into the model, it stores attention states in memory. A 32k context window on a 70B model requires several extra Gigabytes of RAM independent of the model weights.
  • Overhead: This estimate reserves 1.5 GB for the runtime. Actual allocation depends on the device, framework and workload, so compare against usable rather than advertised memory.