Datacenter GPUNVIDIA

NVIDIA H100 SXM5 80GB

Vendor Claim·2025-01-0513 workloads · 21 records
MyAI Rating9.4emb/s · Embedding throughput

VRAM

80 GB

TDP

700 W

MSRP

$25k

Perf/W

0.40 emb/s/W

Cost/1K tok

$0.94/M

Tested

2025-01-05

Quick answer

How many tokens per second does NVIDIA H100 SXM5 80GB produce on Llama 3 70B Q4?

NVIDIA H100 SXM5 80GB produces approximately 66.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 80GB VRAM, 700W TDP). That figure comes from 5 measured runs on TensorRT-LLM 0.10. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-h100-l3-70b-q4)As of 2024-06-24

Overview

Datacenter Workhorse

NVIDIA H100 SXM5 80GB — Hopper-architecture datacenter GPU. 80GB HBM3 at 3.35 TB/s, 700W TDP, with Transformer Engine for FP8 acceleration. NVLink and NVSwitch for scale-out. The reference platform for production LLM serving through 2025.

AI Usefulness

Comfortably hosts 70B FP16 with substantial context. 405B at Q4 fits with headroom. Production-grade throughput at batch sizes 8-64 via TensorRT-LLM or vLLM. The default datacenter inference GPU — used by every major cloud provider. Overkill for single-user workloads; right answer for multi-tenant serving and training runs.

Editorial Verdict

9.8
Editor Rating

The default datacenter inference GPU through 2025. Right answer for production multi-tenant serving. Not for homelab — this is enterprise/cloud territory.

What it does well

  • +80GB HBM3 at 3.35 TB/s — production-grade serving for 70B FP16
  • +Transformer Engine with FP8 acceleration
  • +NVLink + NVSwitch for scale-out to 8-GPU DGX configurations
  • +The reference platform for every major cloud provider's GPU instances

Where it breaks

  • , 700W TDP — datacenter power and cooling required
  • , $25,000-30,000 per card — not consumer-accessible
  • , Overkill for single-user workloads
  • , Being superseded by H200 for capacity and B200 for compute

Sweet Spot

70B FP16 with substantial context, 405B Q4 with headroom. Production multi-tenant serving at batch 8-64 via TensorRT-LLM or vLLM.

Bad Use Cases

  • ×Single-user homelab (absurd overkill)
  • ×Budget-constrained deployments (A100 is more accessible)
  • ×Edge or embedded (physically impossible)
  • ×Anyone who has to ask the price

What Breaks First

Thermal throttling at sustained 700W without proper datacenter cooling. NVLink fabric errors under heavy multi-GPU load if interconnect cables aren't properly seated.

Software Support

vLLMTensorRT-LLMSGLangPyTorchDeepSpeedNVIDIA Tritonllama.cpp (CUDA)

Ubuntu 22.04/24.04 LTS (reference), RHEL 9, Windows Server (limited)

Best Pairings

  • TensorRT-LLM + Llama 3.1 70B FP8 for max throughput
  • vLLM + AWQ-INT4 70B for multi-tenant serving
  • DGX H100 8-GPU configuration for 405B-scale inference

Power & Cooling

700W SXM. Requires datacenter power infrastructure. Not suitable for residential electrical. Paired with 2kW+ per node when accounting for CPU, memory, and cooling overhead.

Verdict

NVIDIA H100 SXM5 80GB with 80GB VRAM at 700W TDP, scored across 13 workloads with 21 benchmark records.

Best workload

Embedding throughput

12400 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0150300450600Llama 38B FP16Llama 370B Q4Llama 370B Q8DeepSeek-R17BGemma 29BQwen 2.514B

Benchmarks (13 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B FP16

llm

282.0tok/sFP164K±11.3Vendor Claim2025-01-05
Llama 3 70B Q4

llm

410.0tok/sQ4_K_M4K, Curated Aggregate2024-08-30
Llama 3 70B Q8

llm

42.0tok/sQ8_04K, Curated Aggregate2024-06-20
SDXL image gen

image

28.0img/minFP16, , Curated Aggregate2026-04-09
Whisper transcription

audio

145.0x RTFP16, , Curated Aggregate2024-08-16
Embedding throughput

embedding

12400.0emb/sFP161K, Curated Aggregate2025-01-12
DeepSeek-R1 7B

llm

285.0tok/sQ4_K_M16K, Curated Aggregate2025-04-25
Gemma 2 9B

llm

235.0tok/sQ4_K_M8K, Curated Aggregate2025-05-22
Qwen 2.5 14B

llm

168.0tok/sQ4_K_M8K, Curated Aggregate2024-10-21
Mistral 7B

llm

295.0tok/sQ4_K_M16K, Curated Aggregate2025-05-08
Phi-3 Mini

llm

612.0tok/sQ4_K_M4K, Curated Aggregate2025-06-04
Llama 3 8B Q4

llm

178.0tok/sQ4_K_M128K, Curated Aggregate2024-11-08
Asure 12B

llm

195.0tok/sQ4_K_M8K±7.8Vendor Claim2025-05-01

MyAI Score

Reference
9.4/10

NVIDIA H100 SXM5 80GB clears a 9.4/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
354
Capability
210
Efficiency
55
Value
27
Trust
45
Coverage
60
Composite benchmark751 / 1000

Workload Fit

What models fit this 80GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Good
70B
Excellent
Good
Won't fit
Top Benchmarks
Embedding throughput12400 emb/s
Phi-3 Mini612 tok/s
Llama 3 70B Q4410 tok/s

Source

Text-embeddings-inference, batch 32.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.4

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2025-01-12; 592 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $25k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.