How many tokens per second does AMD Instinct MI300X 192GB produce on Llama 3 70B Q4?
AMD Instinct MI300X 192GB produces approximately 72.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 192GB VRAM, 750W TDP). That figure comes from 5 measured runs on vLLM 0.6.3 ROCm 6.2. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.