Interactive Utility
DeepSeek V3 (671B) VRAM Calculator
Stop guessing your hardware needs for DeepSeek V3 (671B). Calculate the exact VRAM required for inference based on parameter count, quantization, and KV cache, and see which machines can actually run it.
Model
Custom: billion
The community default. Best size-to-quality ratio.
Memory to run it
- Weights (Q4_K_M)
- 406.8 GB
- KV cache (8k)
- 8.0 GB
- Runtime overhead
- 1.5 GB
- System RAM to load
- 409 GB
- Model file on disk
- 427.1 GB
Hardware match
Multi-GPU node or Mac Studio (192–512GB)
A dedicated server. Below this, a hosted API is usually cheaper.
Rough decode speed
~4 tok/s
Ballpark on H100 / multi-GPU, bounded by memory bandwidth. Real numbers land 30–50% either side depending on framework, prompt length and thermals.
DeepSeek V3 (671B) memory requirements by quantisation
Estimated at an 8k-token context. DeepSeek V3 (671B) is 671B parameters; the arithmetic and its assumptions are set out below.
| Quantisation | Weights | Total VRAM | Runs on |
|---|---|---|---|
| FP16 (unquantised) | 1342.0 GB | 1351.5 GB | Multi-GPU node or Mac Studio (192–512GB) |
| Q8_0 (8-bit) | 712.9 GB | 722.4 GB | Multi-GPU node or Mac Studio (192–512GB) |
| Q6_K (6-bit) | 553.6 GB | 563.1 GB | Multi-GPU node or Mac Studio (192–512GB) |
| Q5_K_M (5-bit) | 478.1 GB | 487.6 GB | Multi-GPU node or Mac Studio (192–512GB) |
| Q4_K_M (4-bit) | 406.8 GB | 416.3 GB | Multi-GPU node or Mac Studio (192–512GB) |
| Q3_K_M (3-bit) | 327.1 GB | 336.6 GB | Multi-GPU node or Mac Studio (192–512GB) |
Arithmetic over stated assumptions, not a benchmark. Weights are parameter count times bits-per-weight; the KV cache term assumes grouped-query attention; 1.5GB is reserved for runtime overhead.
Size a related model
Embed this DeepSeek V3 (671B) Calculator
Add this exact configuration to your blog, documentation, or internal wiki.
<iframe src="https://fathomlayer.com/embed/hardware-calculator?p=671" width="100%" height="600" frameborder="0" style="border-radius: 12px; border: 1px solid rgba(255,255,255,0.1);"></iframe>
How this math works
Running Large Language Models locally is entirely bottlenecked by memory. Raw compute (TFLOPS) dictates your generation speed, but VRAM capacity dictates if the model will load at all.
- Model Weights: At FP16 (unquantized), every 1 Billion parameters requires ~2GB of VRAM. At Q4 (4-bit quantization), that drops to ~0.7GB per 1B parameters.
- KV Cache (Context Window): As you feed text into the model, it stores attention states in memory. A 32k context window on a 70B model requires several extra Gigabytes of RAM independent of the model weights.
- Overhead: CUDA and operating systems reserve memory (usually 1-2GB), meaning a 24GB RTX 4090 cannot realistically load a 23.5GB model.
