How many tokens per second does NVIDIA H100 SXM5 80GB produce on Llama 3 70B Q4?
NVIDIA H100 SXM5 80GB produces approximately 66.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 80GB VRAM, 700W TDP). That figure comes from 5 measured runs on TensorRT-LLM 0.10. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.