Datacenter GPUAMD

AMD Instinct MI250X 128GB

Curated Aggregate·2024-11-124 workloads · 4 records
MyAI Rating8.5tok/s · Llama 3 8B FP16

VRAM

128 GB

TDP

560 W

MSRP

$14k

Perf/W

0.09 tok/s/W

Cost/1K tok

$0.0028/k

Tested

2024-11-12

Quick answer

How many tokens per second does AMD Instinct MI250X 128GB produce on Llama 3 70B Q4?

AMD Instinct MI250X 128GB produces approximately 52.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 128GB VRAM, 560W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-mi250x-l3-70b-q4)As of 2024-11-12

Overview

AMD Datacenter Legacy

AMD Instinct MI250X 128GB — CDNA2 datacenter accelerator. Dual-die design with 128GB HBM2e at 3.2 TB/s, 560W TDP. AMD's pre-MI300 datacenter GPU with mature ROCm support.

AI Usefulness

128GB VRAM at 3.2 TB/s bandwidth — competitive with H100 for memory-bound inference. ROCm 6.x support is production-grade for vLLM and llama.cpp. ~110 tok/s on 12B Q4. Best for: AMD-friendly datacenter deployments, large-model inference (70B+ FP16), and organizations with existing ROCm infrastructure.

Verdict

AMD Instinct MI250X 128GB with 128GB VRAM at 560W TDP, scored across 4 workloads with 4 benchmark records.

Best workload

Llama 3 8B FP16

188 tok/s

Quantization

FP16

4K context · batch 1

LLM Inference Performance

050100150200Llama 370B Q4Llama 38B FP16Mistral 7BGemma 29B

Benchmarks (4 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

52.0tok/sQ4_K_M8K, Curated Aggregate2024-11-12
Llama 3 8B FP16

llm

188.0tok/sFP164K, Curated Aggregate2024-11-19
Mistral 7B

llm

135.0tok/sQ4_K_M4K, Curated Aggregate2024-04-08
Gemma 2 9B

llm

110.0tok/sQ4_K_M8K, Curated Aggregate2024-07-31

MyAI Score

Flagship
8.5/10

AMD Instinct MI250X 128GB clears a 8.5/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
277
Capability
185
Efficiency
37
Value
31
Trust
43
Coverage
25
Composite benchmark598 / 1000

Workload Fit

What models fit this 128GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Tight
Top Benchmarks
Llama 3 8B FP16188 tok/s
Mistral 7B135 tok/s
Gemma 2 9B110 tok/s

Public Trust Layer

Trust score

6/10

MyAI rating

8.5

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-11-19; 646 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $14k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is a strong buy when the listing stays close to reference pricing and matches the workload you actually run.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.