What hardware do I need to run Llama 3.3 70B?
Llama 3.3 70B needs at minimum 42.4 GB of VRAM at Q4_K_M quantization (154 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4 Pro, 64 GB) (64 GB VRAM, $2,199 MSRP). Measured throughput hits 540 tok/s on NVIDIA H200 141GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.