Consumer GPUAMD

AMD Radeon RX 7900 XTX 24GB

Curated Aggregate·2025-06-0912 workloads · 19 records
MyAI Rating9.0emb/s · Embedding throughput

VRAM

24 GB

TDP

355 W

MSRP

$999

Perf/W

0.27 emb/s/W

Cost/1K tok

$0.11/M

Tested

2025-06-09

Quick answer

How many tokens per second does AMD Radeon RX 7900 XTX 24GB produce on Llama 3 70B Q4?

AMD Radeon RX 7900 XTX 24GB produces approximately 11.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 24GB VRAM, 355W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-rx7900xtx-l3-70b-q4)As of 2025-08-13

Overview

AMD Value

AMD Radeon RX 7900 XTX 24GB — RDNA 3 consumer flagship. 24GB GDDR6 at ~960 GB/s, 355W TDP. Competes with RTX 4080/4090 on VRAM at dramatically lower price. ROCm support improving but not at CUDA parity.

AI Usefulness

24GB VRAM at ~$900-1000 makes this the best $/VRAM consumer card if you can tolerate the ROCm ecosystem. Llama 3 8B Q4 runs at ~90+ tok/s via llama.cpp ROCm. Vulkan backends (llama.cpp) are slower but work universally. Best for Linux-native builders who want NVIDIA-tier VRAM at AMD pricing and are comfortable with the software tradeoffs. Not recommended for Windows AI workflows or production CUDA-dependent stacks.

Editorial Verdict

6.5
Editor Rating

Buy for Linux AI builds where $/VRAM is the priority and you're comfortable with the ROCm ecosystem. Skip for Windows, for production, or if you value plug-and-play over tinkering.

What it does well

  • +24GB GDDR6 at ~960 GB/s — NVIDIA-competitive VRAM at ~$900-1000
  • +Best $/VRAM consumer card if you tolerate ROCm ecosystem
  • +355W TDP — manageable with quality 850W PSU
  • +llama.cpp Vulkan backend works universally without ROCm

Where it breaks

  • , ROCm software stack trails CUDA in maturity and breadth
  • , Vulkan paths are slower than native ROCm
  • , $900-1000 competes with used 3090 which has CUDA
  • , No FP8 acceleration — missing modern quantization features
  • , Not recommended for Windows AI workflows

Sweet Spot

24GB VRAM for Linux-native builders who want NVIDIA-tier capacity at AMD pricing and are comfortable with software tradeoffs

Bad Use Cases

  • ×Windows AI (ROCm is Linux-only, Vulkan is slow)
  • ×Production CUDA-dependent stacks
  • ×Beginners who want plug-and-play (NVIDIA is the answer)
  • ×vLLM at scale (ROCm support is improving but not at CUDA parity)

What Breaks First

ROCm driver compatibility — pin your kernel and ROCm versions once stable. HIP runtime errors are more common than CUDA errors in edge cases. Vulkan fallback is functional but 30-50% slower.

Software Support

llama.cpp (ROCm/HIP, Vulkan)Ollama (ROCm experimental)vLLM (ROCm, maturing)PyTorch (ROCm)

Ubuntu 22.04/24.04 (good with ROCm), Windows (Vulkan only, slower), macOS (unsupported)

Best Pairings

  • llama.cpp ROCm + Llama 3.1 8B Q4 for 90+ tok/s
  • Ubuntu 22.04 + ROCm 6.2.4 — the validated stack
  • 850W Gold PSU

Power & Cooling

355W TGP. 850W PSU minimum. ROCm power management less mature than NVIDIA. Higher idle power draw than equivalent NVIDIA cards.

Verdict

AMD Radeon RX 7900 XTX 24GB with 24GB VRAM at 355W TDP, scored across 12 workloads with 19 benchmark records.

Best workload

Embedding throughput

5800 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

050100150200Llama 38B Q4Llama 370B Q4Llama 38B FP16Qwen 2.514BMistral 7BPhi-3 Mini

Benchmarks (12 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

95.0tok/sQ4_K_M4K, Curated Aggregate2025-06-09
Llama 3 70B Q4

llm

24.0tok/sQ4_K_M4K, Curated Aggregate2025-08-06
SDXL image gen

image

13.0img/minFP16, , Curated Aggregate2026-03-11
Llama 3 8B FP16

llm

88.0tok/sFP164K, Curated Aggregate2024-05-18
Qwen 2.5 14B

llm

52.0tok/sQ4_K_M8K, Curated Aggregate2024-10-11
Whisper transcription

audio

48.0x RTFP16, , Curated Aggregate2024-09-30
Mistral 7B

llm

95.0tok/sQ4_K_M4K, Curated Aggregate2025-02-12
Phi-3 Mini

llm

195.0tok/sQ4_K_M4K, Curated Aggregate2025-03-08
Gemma 2 9B

llm

72.0tok/sQ4_K_M8K, Curated Aggregate2024-09-05
DeepSeek-R1 7B

llm

78.0tok/sQ4_K_M8K, Curated Aggregate2025-04-15
Embedding throughput

embedding

5800.0emb/sFP161K, Curated Aggregate2024-11-04
Asure 12B

llm

52.0tok/sQ4_K_M8K±3.6Curated Aggregate2025-05-01

MyAI Score

Reference
9.0/10

AMD Radeon RX 7900 XTX 24GB clears a 9.0/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
282
Capability
159
Efficiency
45
Value
55
Trust
45
Coverage
60
Composite benchmark646 / 1000

Workload Fit

What models fit this 24GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Tight
32B
Excellent
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Top Benchmarks
Embedding throughput5800 emb/s
Phi-3 Mini195 tok/s
Llama 3 8B Q495 tok/s

Source

TEI on ROCm 6.2.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.0

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-11-04; 661 days old.

Open primary source

Verified Local-Lab Run

Measured on your own machine via ollama-local-api on 2026-06-01.

Windows 10 build 26200 + ROCm 7.10 + Ollama 0.24.0 (HIPBLAS). 4 runs per model after a discarded warmup; decode_tps reported as per-run eval_count/eval_duration. Variance <= 0.4% on reproduced models. Captured by site owner on declared rig. Prompt SHA-256: 5e752b688e0761f321a5721d2fa6b57678bf1de1a1180bb4fc9b7e709bafb6a5.

llama-3.2-3b-instruct

3B / Q4_K_M

202.10 tok/s

median 201.87 / min 201.31 / max 202.80 / n=4

qwen-2.5-7b-instruct

7B / Q4_K_M

115.60 tok/s

median 115.44 / min 115.27 / max 116.04 / n=4

mistral-7b-instruct

7B / Q4_K_M

118.79 tok/s

median 118.72 / min 118.38 / max 119.28 / n=4

gemma-4-31b-it

31B / Q4_K_M

decode failed

n=4, throughput 0 tok/s

Decode failed: model prefilled then produced 0 tokens across all runs on AMD ROCm 7.10 + Ollama 0.24.0. A real, reproducible failure mode, not missing data.

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $999 as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.