Compute
Local inference — reference figures
Indicative throughput, power and hourly cost for running open-weight models locally. Use them to rank options and sanity-check a build — not as a spec sheet.
Normalised toBatch 1 · ~1k prompt
SourcePublic benchmarks
| Hardware | Model / quant | Throughput | Power | $/hr | Framework |
|---|---|---|---|---|---|
| NVIDIA H100 (80GB SXM5)Cloud rental rate | Llama-3-70BFP16 | 285.0tok/s | 700W | $2.89 | TensorRT-LLM |
| Mac Studio M2 Ultra (192GB)Memory-bandwidth bound | Llama-3-70BQ4_K_M | 18.5tok/s | 130W | $0.12 | MLX |
| 2x RTX 4090 (48GB)Tensor-parallel over PCIe | Llama-3-70BEXL2 4.0bpw | 34.0tok/s | 850W | $0.25 | ExLlamaV2 |
| RTX 4090 (24GB) | Llama-3-8BFP16 | 145.0tok/s | 400W | $0.15 | vLLM |
| MacBook Pro M3 Max (128GB)Sustained draw throttles | Mixtral 8x7BQ4_0 | 27.0tok/s | 65W | $0.08 | llama.cpp |
| 2x Tesla P40 (48GB)No flash-attention; used-market rig | Llama-3-70BQ4_K_M | 9.8tok/s | 500W | $0.09 | llama.cpp |
| AMD RX 7900 XTX (24GB) | Mistral-7BQ8_0 | 42.0tok/s | 355W | $0.12 | llama.cpp / ROCm |
Figures are compiled from published community benchmarks and vendor documentation and normalised to a single request with a ~1k-token prompt. Real numbers move with driver version, prompt length, batching and thermal headroom — a ±20% band is normal. To size a specific model, use the hardware sizer.
