Datacenter GPUNVIDIA

NVIDIA GH200 Grace Hopper 480GB

Curated Aggregate·2025-02-184 workloads · 4 records
MyAI Rating8.9tok/s · Llama 3 8B FP16

VRAM

96 GB

TDP

1000 W

MSRP

$44k

Perf/W

0.10 tok/s/W

Cost/1K tok

$0.0049/k

Tested

2025-02-18

Quick answer

How many tokens per second does NVIDIA GH200 Grace Hopper 480GB produce on Llama 3 70B Q4?

NVIDIA GH200 Grace Hopper 480GB produces approximately 95.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 96GB VRAM, 1000W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-gh200-l3-70b-q4)As of 2025-02-18

Overview

Superchip Frontier

NVIDIA GH200 Grace Hopper Superchip — combines 72-core Grace ARM CPU with H100 GPU and 480GB LPDDR5X unified memory. 144GB HBM3e GPU memory, 1000W TDP. NVLink-C2C interconnect.

AI Usefulness

624GB total accessible memory (144GB HBM3e + 480GB LPDDR5X via NVLink-C2C) makes this the most memory-rich single-socket AI platform. Hosts Llama 3.1 405B at Q4 with massive context. ~248 tok/s on 12B Q4. Best for: frontier model inference, training up to 70B, and memory-bandwidth-bound HPC workloads. Overkill for sub-70B inference.

Verdict

NVIDIA GH200 Grace Hopper 480GB with 96GB VRAM at 1000W TDP, scored across 4 workloads with 4 benchmark records.

Best workload

Llama 3 8B FP16

380 tok/s

Quantization

FP16

8K context · batch 1

LLM Inference Performance

095190285380Llama 370B Q4Llama 38B FP16Mistral 7BGemma 29B

Benchmarks (4 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

95.0tok/sQ4_K_M8K, Curated Aggregate2025-02-18
Llama 3 8B FP16

llm

380.0tok/sFP168K, Curated Aggregate2025-02-25
Mistral 7B

llm

305.0tok/sQ4_K_M4K, Curated Aggregate2024-08-04
Gemma 2 9B

llm

248.0tok/sQ4_K_M8K, Curated Aggregate2024-10-26

MyAI Score

Flagship
8.9/10

NVIDIA GH200 Grace Hopper 480GB clears a 8.9/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
320
Capability
185
Efficiency
40
Value
24
Trust
43
Coverage
25
Composite benchmark637 / 1000

Workload Fit

What models fit this 96GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Won't fit
Top Benchmarks
Llama 3 8B FP16380 tok/s
Mistral 7B305 tok/s
Gemma 2 9B248 tok/s

Source

Vendor benchmark.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

8.9

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2025-02-25; 548 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $44k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is a strong buy when the listing stays close to reference pricing and matches the workload you actually run.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.