Skip to content

Compute

Build blueprints

Reference parts lists for local AI machines — the components to hit a target model and a throughput band, priced from what you can buy now. Estimates, not test rigs; confirm the model fit with the hardware sizer.

Blueprints

The 70B quantised workhorse

Llama-3-70B (Q4_K_M), roughly 12-16 tok/s

~$6,500est. build cost
  • GPU2x NVIDIA RTX 4090 (24GB)
  • CPUAMD Ryzen Threadripper 7960X
  • RAM128GB DDR5 ECC (4x32GB)
  • PSU1600W Titanium, ATX 3.0

48GB of VRAM fits a 4-bit 70B with room for context. The bottleneck is the PCIe link between the two cards — expect lower throughput than a single card with the same VRAM.

The datacenter node

Mixtral 8x22B (FP16), roughly 20-30 tok/s

~$22,000est. build cost
  • GPU4x NVIDIA RTX 6000 Ada (48GB)
  • CPUAMD EPYC 9354P (32 core)
  • RAM256GB DDR5 ECC RDIMM
  • PSU2x 1600W redundant Titanium

192GB of VRAM runs a full-precision Mixtral or several quantised 70Bs concurrently. Blower-style cards and real airflow are not optional at this density.

Reader setups

No reader setups published yet. Send yours — parts list, the model you run and the throughput you measured — through the contact page.