The 70B quantised workhorse
Llama-3-70B (Q4_K_M), roughly 12-16 tok/s
- GPU2x NVIDIA RTX 4090 (24GB)
- CPUAMD Ryzen Threadripper 7960X
- RAM128GB DDR5 ECC (4x32GB)
- PSU1600W Titanium, ATX 3.0
48GB of VRAM fits a 4-bit 70B with room for context. The bottleneck is the PCIe link between the two cards — expect lower throughput than a single card with the same VRAM.
