What hardware do I need to run Llama 3.1 405B?
Llama 3.1 405B needs at minimum 245 GB of VRAM at Q4_K_M quantization (891 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac Studio (M3 Ultra, 512 GB) (512 GB VRAM, $9,499 MSRP). Community benchmark submissions are open. This model exceeds 48 GB at Q4, so plan for a 2-or-more-GPU split.