Datacenter GPUNVIDIA

NVIDIA H200 141GB

Vendor Claim·2024-04-0611 workloads · 14 records
MyAI Rating9.4emb/s · Embedding throughput

VRAM

141 GB

TDP

700 W

MSRP

$30k

Perf/W

0.12 emb/s/W

Cost/1K tok

$0.0038/k

Tested

2024-04-06

Quick answer

How many tokens per second does NVIDIA H200 141GB produce on Llama 3 70B Q4?

NVIDIA H200 141GB produces approximately 84.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 141GB VRAM, 700W TDP). That figure comes from 5 measured runs on TensorRT-LLM 0.14. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-h200-l3-70b-q4)As of 2024-04-06

Overview

Capacity King

NVIDIA H200 141GB — Hopper-generation capacity-upgrade GPU. 141GB HBM3e at 4.8 TB/s, 700W TDP. Same compute dies as H100 with dramatically more and faster memory. Purpose-built for large-model inference.

AI Usefulness

141GB VRAM hosts Llama 3.1 405B at FP8 comfortably. The go-to for frontier-model serving in 2025-2026. ~2x the memory bandwidth of H100 means ~1.8x tok/s on memory-bound decode. The capacity play — not a compute upgrade over H100, but the VRAM ceiling boost is what matters for the largest open models.

Editorial Verdict

9.5
Editor Rating

The capacity king. Right answer when 80GB isn't enough and you need 141GB. For everything else, H100 or L40S are more economical.

What it does well

  • +141GB HBM3e at 4.8 TB/s — 2× the memory bandwidth of H100
  • +Hosts Llama 3.1 405B at FP8 comfortably
  • +Same Hopper compute dies as H100 — software compatibility is identical
  • +The go-to for frontier-model serving in 2025-2026

Where it breaks

  • , 700W TDP — datacenter only
  • , $30,000+ — enterprise pricing
  • , Not a compute upgrade over H100 — it's a capacity play
  • , Supply constrained — cloud availability limited

Sweet Spot

405B-class models at FP8 with context headroom — the capacity play for frontier open-weight models

Bad Use Cases

  • ×Workloads that fit in 80GB (H100 is cheaper, same compute)
  • ×Budget-constrained deployments
  • ×Homelab (physically and financially impossible)

What Breaks First

HBM3e thermal management at sustained 4.8 TB/s bandwidth — requires proper datacenter airflow. Memory errors under thermal stress more common than H100 due to higher density.

Software Support

vLLMTensorRT-LLMSGLangPyTorchDeepSpeedNVIDIA Triton

Ubuntu 22.04/24.04 LTS (reference), RHEL 9

Best Pairings

  • vLLM + Llama 3.1 405B AWQ-INT4 for maximum capacity utilization
  • TensorRT-LLM FP8 path for frontier throughput

Power & Cooling

700W SXM. Datacenter infrastructure required. The 141GB HBM3e draws more board power than H100 under sustained memory load.

Verdict

NVIDIA H200 141GB with 141GB VRAM at 700W TDP, scored across 11 workloads with 14 benchmark records.

Best workload

Embedding throughput

15200 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0150300450600Llama 370B Q4Llama 370B Q8Llama 38B FP16Mistral 7BDeepSeek-R17BGemma 29B

Benchmarks (11 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

540.0tok/sQ4_K_M4K, Curated Aggregate2024-11-22
Llama 3 70B Q8

llm

58.0tok/sQ8_08K, Curated Aggregate2024-04-10
Llama 3 8B FP16

llm

340.0tok/sFP164K, Curated Aggregate2025-11-18
Mistral 7B

llm

305.0tok/sQ4_K_M8K, Curated Aggregate2026-01-28
SDXL image gen

image

36.0img/minFP16, , Curated Aggregate2026-02-15
DeepSeek-R1 7B

llm

305.0tok/sQ4_K_M8K, Curated Aggregate2026-04-20
Gemma 2 9B

llm

220.0tok/sQ4_K_M8K, Curated Aggregate2024-07-30
Qwen 2.5 14B

llm

178.0tok/sQ4_K_M8K, Curated Aggregate2024-12-15
Llama 3 8B Q4

llm

215.0tok/sQ4_K_M128K, Curated Aggregate2025-02-04
Embedding throughput

embedding

15200.0emb/sFP161K, Curated Aggregate2024-09-08
Asure 12B

llm

230.0tok/sQ4_K_M8K±9.2Vendor Claim2025-05-01

MyAI Score

Reference
9.4/10

NVIDIA H200 141GB clears a 9.4/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
357
Capability
210
Efficiency
57
Value
30
Trust
48
Coverage
55
Composite benchmark757 / 1000

Workload Fit

What models fit this 141GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Tight
Top Benchmarks
Embedding throughput15200 emb/s
Llama 3 70B Q4540 tok/s
Llama 3 8B FP16340 tok/s

Source

TEI 1.5 on H200.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.4

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-09-08; 718 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $30k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.