Skip to content

Compute

Local inference — reference figures

Indicative throughput, power and hourly cost for running open-weight models locally. Use them to rank options and sanity-check a build — not as a spec sheet.

Normalised toBatch 1 · ~1k prompt
SourcePublic benchmarks
HardwareModel / quantThroughputPower$/hrFramework
NVIDIA H100 (80GB SXM5)Cloud rental rateLlama-3-70BFP16285.0tok/s700W$2.89TensorRT-LLM
Mac Studio M2 Ultra (192GB)Memory-bandwidth boundLlama-3-70BQ4_K_M18.5tok/s130W$0.12MLX
2x RTX 4090 (48GB)Tensor-parallel over PCIeLlama-3-70BEXL2 4.0bpw34.0tok/s850W$0.25ExLlamaV2
RTX 4090 (24GB)Llama-3-8BFP16145.0tok/s400W$0.15vLLM
MacBook Pro M3 Max (128GB)Sustained draw throttlesMixtral 8x7BQ4_027.0tok/s65W$0.08llama.cpp
2x Tesla P40 (48GB)No flash-attention; used-market rigLlama-3-70BQ4_K_M9.8tok/s500W$0.09llama.cpp
AMD RX 7900 XTX (24GB)Mistral-7BQ8_042.0tok/s355W$0.12llama.cpp / ROCm

Figures are compiled from published community benchmarks and vendor documentation and normalised to a single request with a ~1k-token prompt. Real numbers move with driver version, prompt length, batching and thermal headroom — a ±20% band is normal. To size a specific model, use the hardware sizer.