Datacenter GPUAMD

AMD Instinct MI300X 192GB

Curated Aggregate·2025-07-1112 workloads · 20 records
MyAI Rating9.3emb/s · Embedding throughput

VRAM

192 GB

TDP

750 W

MSRP

$18k

Perf/W

0.10 emb/s/W

Cost/1K tok

$0.0026/k

Tested

2025-07-11

Quick answer

How many tokens per second does AMD Instinct MI300X 192GB produce on Llama 3 70B Q4?

AMD Instinct MI300X 192GB produces approximately 72.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 192GB VRAM, 750W TDP). That figure comes from 5 measured runs on vLLM 0.6.3 ROCm 6.2. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-mi300x-l3-70b-q4)As of 2025-07-11

Overview

AMD Datacenter

AMD Instinct MI300X 192GB — CDNA3-architecture datacenter APU. 192GB HBM3 at 5.3 TB/s, 750W TDP. AMD's answer to the H100/H200 for AI workloads. ROCm software stack maturing through 2025-2026.

AI Usefulness

192GB VRAM at competitive bandwidth makes this a credible H200 alternative for inference. vLLM ROCm support is production-grade. Llama 3.1 70B Q4 runs at ~120+ tok/s. The software tax (ROCm vs CUDA) is the real consideration — fewer engines support it, and community tooling is thinner. Best for operators who already have AMD infrastructure or who are specifically chasing $/VRAM ratios.

Editorial Verdict

8
Editor Rating

Best for organizations with existing AMD infrastructure or those specifically chasing $/VRAM ratios at datacenter scale. The software maturity gap vs NVIDIA is real but narrowing.

What it does well

  • +192GB HBM3 at 5.3 TB/s — the highest VRAM single-GPU available
  • +ROCm 6.x is production-grade for vLLM and llama.cpp
  • +Competes with H200 on capacity at potentially lower cost
  • +~120+ tok/s on 70B Q4 — competitive with NVIDIA datacenter

Where it breaks

  • , Software tax: ROCm ecosystem trails CUDA in breadth and community tooling
  • , Fewer engines support it — no TensorRT-LLM, no ExLlamaV2
  • , 750W TDP — datacenter power and cooling required
  • , $10,000-15,000 — not consumer territory

Sweet Spot

Large-model inference where VRAM capacity matters more than peak compute — 405B Q4 with headroom, 70B FP16 comfortably

Bad Use Cases

  • ×CUDA-dependent workflows
  • ×Consumer/homelab builds
  • ×Windows AI (ROCm is Linux-only)
  • ×Workloads requiring TensorRT-LLM optimizations

What Breaks First

ROCm driver compatibility with specific kernel versions — pin your ROCm + kernel combination once stable. vLLM ROCm branch occasionally lags upstream.

Software Support

vLLM (ROCm)llama.cpp (ROCm/HIP)PyTorch (ROCm)Ollama (ROCm experimental)

Ubuntu 22.04 LTS (reference), RHEL 9, other Linux (ROCm-supported), Windows (unsupported), macOS (unsupported)

Best Pairings

  • vLLM 0.6.3+ ROCm 6.2 + Llama 3.1 70B AWQ-INT4 for serving
  • llama.cpp HIP + 70B Q4 for single-stream
  • Ubuntu 22.04 + ROCm 6.2.4 — the validated stack

Power & Cooling

750W TDP. Datacenter power and cooling required. Not suitable for residential. ROCm power management less mature than NVIDIA's — expect higher idle draw.

Verdict

AMD Instinct MI300X 192GB with 192GB VRAM at 750W TDP, scored across 12 workloads with 20 benchmark records.

Best workload

Embedding throughput

9800 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

095190285380Llama 370B Q4Llama 370B Q8Llama 38B FP16Qwen 2.514BDeepSeek-R17BGemma 29B

Benchmarks (12 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

380.0tok/sQ4_K_M4K, Curated Aggregate2024-09-12
Llama 3 70B Q8

llm

48.0tok/sQ8_08K, Curated Aggregate2025-07-15
Llama 3 8B FP16

llm

260.0tok/sFP168K, Curated Aggregate2025-07-30
SDXL image gen

image

22.0img/minFP16, , Curated Aggregate2024-12-22
Embedding throughput

embedding

9800.0emb/sFP161K, Curated Aggregate2024-11-30
Qwen 2.5 14B

llm

142.0tok/sQ4_K_M8K, Curated Aggregate2024-12-04
DeepSeek-R1 7B

llm

215.0tok/sQ4_K_M8K, Curated Aggregate2026-03-19
Gemma 2 9B

llm

188.0tok/sQ4_K_M8K, Curated Aggregate2025-12-08
Mistral 7B

llm

268.0tok/sQ4_K_M4K, Curated Aggregate2025-05-19
Llama 3 8B Q4

llm

920.0tok/sQ4_K_M4K, Curated Aggregate2024-12-08
Whisper transcription

audio

95.0x RTFP16, , Curated Aggregate2025-01-12
Asure 12B

llm

130.0tok/sQ4_K_M8K±6.5Curated Aggregate2025-05-01

MyAI Score

Reference
9.3/10

AMD Instinct MI300X 192GB clears a 9.3/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
341
Capability
202
Efficiency
50
Value
34
Trust
46
Coverage
60
Composite benchmark733 / 1000

Workload Fit

What models fit this 192GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Excellent
Top Benchmarks
Embedding throughput9800 emb/s
Llama 3 8B Q4920 tok/s
Llama 3 70B Q4380 tok/s

Source

TEI ROCm MI300X.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.3

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-11-30; 635 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $18k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.