Consumer GPUAMD

AMD Radeon RX 7900 XTX 24GB

Curated Aggregate·2025-06-0911 workloads · 17 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

24 GB

TDP

355 W

MSRP

$999

Perf/W

0.27 img/min/W

Cost/1K tok

$0.11/M

Tested

2025-06-09

Quick answer

How many tokens per second does AMD Radeon RX 7900 XTX 24GB produce on Llama 3 70B Q4?

AMD Radeon RX 7900 XTX 24GB has a source-attributed result of 11.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-rx7900xtx-l3-70b-q4)As of 2025-08-13

Overview

AMD Value

AMD Radeon RX 7900 XTX 24GB — RDNA 3 consumer flagship. 24GB GDDR6 at ~960 GB/s, 355W TDP. Competes with RTX 4080/4090 on VRAM at dramatically lower price. ROCm support improving but not at CUDA parity.

AI Usefulness

24GB VRAM at ~$900-1000 makes this the best $/VRAM consumer card if you can tolerate the ROCm ecosystem. Llama 3 8B Q4 runs at ~90+ tok/s via llama.cpp ROCm. Vulkan backends (llama.cpp) are slower but work universally. Best for Linux-native builders who want NVIDIA-tier VRAM at AMD pricing and are comfortable with the software tradeoffs. Not recommended for Windows AI workflows or production CUDA-dependent stacks.

Editorial Verdict

6.5
Editor Rating

Buy for Linux AI builds where $/VRAM is the priority and you're comfortable with the ROCm ecosystem. Skip for Windows, for production, or if you value plug-and-play over tinkering.

What it does well

  • +24GB GDDR6 at ~960 GB/s — NVIDIA-competitive VRAM at ~$900-1000
  • +Best $/VRAM consumer card if you tolerate ROCm ecosystem
  • +355W TDP — manageable with quality 850W PSU
  • +llama.cpp Vulkan backend works universally without ROCm

Where it breaks

  • , ROCm software stack trails CUDA in maturity and breadth
  • , Vulkan paths are slower than native ROCm
  • , $900-1000 competes with used 3090 which has CUDA
  • , No FP8 acceleration — missing modern quantization features
  • , Not recommended for Windows AI workflows

Sweet Spot

24GB VRAM for Linux-native builders who want NVIDIA-tier capacity at AMD pricing and are comfortable with software tradeoffs

Bad Use Cases

  • ×Windows AI (ROCm is Linux-only, Vulkan is slow)
  • ×Production CUDA-dependent stacks
  • ×Beginners who want plug-and-play (NVIDIA is the answer)
  • ×vLLM at scale (ROCm support is improving but not at CUDA parity)

What Breaks First

ROCm driver compatibility — pin your kernel and ROCm versions once stable. HIP runtime errors are more common than CUDA errors in edge cases. Vulkan fallback is functional but 30-50% slower.

Software Support

llama.cpp (ROCm/HIP, Vulkan)Ollama (ROCm experimental)vLLM (ROCm, maturing)PyTorch (ROCm)

Ubuntu 22.04/24.04 (good with ROCm), Windows (Vulkan only, slower), macOS (unsupported)

Best Pairings

  • llama.cpp ROCm + Llama 3.1 8B Q4 for 90+ tok/s
  • Ubuntu 22.04 + ROCm 6.2.4 — the validated stack
  • 850W Gold PSU

Power & Cooling

355W TGP. 850W PSU minimum. ROCm power management less mature than NVIDIA. Higher idle power draw than equivalent NVIDIA cards.

Verdict

AMD Radeon RX 7900 XTX 24GB with 24GB VRAM at 355W TDP, scored across 11 workloads with 17 benchmark records.

Reference workload

SDXL image gen

13 img/min

Quantization

FP16

· batch 1

LLM Inference Performance

050100150200Llama 3 8BQ4Llama 370B Q4Qwen 2.514BMistral 7BPhi-3 MiniGemma 29B

Benchmarks (11 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

95.0tok/sQ4_K_M4K, Curated Aggregate2025-06-09
Llama 3 70B Q4

llm

11.0tok/sQ4_K_M4K, Curated Aggregate2025-08-13
SDXL image gen

image

13.0img/minFP16, , Curated Aggregate2026-03-11
Qwen 2.5 14B

llm

48.0tok/sQ4_K_M8K, Curated Aggregate2024-11-27
Whisper transcription

audio

22.0x RTFP16, , Curated Aggregate2024-10-18
Mistral 7B

llm

95.0tok/sQ4_K_M4K, Curated Aggregate2025-02-12
Phi-3 Mini

llm

195.0tok/sQ4_K_M4K, Curated Aggregate2025-03-08
Gemma 2 9B

llm

58.0tok/sQ4_K_M4K, Curated Aggregate2025-03-22
DeepSeek-R1 7B

llm

78.0tok/sQ4_K_M8K, Curated Aggregate2025-04-15
Embedding throughput

embedding

5800.0emb/sFP161K, Curated Aggregate2024-11-04
Asure 12B

llm

52.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score AMD Radeon RX 7900 XTX 24GB. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 24GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Won't fit
32B
Good
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Reported batch-one examples
SDXL image gen13 img/min
Llama 3 70B Q411 tok/s
Llama 3 8B Q495 tok/s

Source

Curated from public sources (llama.cpp logs / vendor specs / vLLM community) — Automatic1111 ROCm.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2026-03-11; 182 days old.

Open primary source

Owner-reported capture

Owner-reported capture via ollama-local-api on 2026-06-01. Hardware identity and outputs are supplied by the owner; this is not independent replication.

Windows 10 build 26200 + ROCm 7.10 + Ollama 0.24.0 (HIPBLAS). 4 runs per model after a discarded warmup; decode_tps reported as per-run eval_count/eval_duration. Variance <= 0.4% on reproduced models. Captured by site owner on declared rig. Prompt SHA-256: 5e752b688e0761f321a5721d2fa6b57678bf1de1a1180bb4fc9b7e709bafb6a5.

llama-3.2-3b-instruct

3B / Q4_K_M

202.10 tok/s

median 201.87 / min 201.31 / max 202.80 / n=4

qwen-2.5-7b-instruct

7B / Q4_K_M

115.60 tok/s

median 115.44 / min 115.27 / max 116.04 / n=4

mistral-7b-instruct

7B / Q4_K_M

118.79 tok/s

median 118.72 / min 118.38 / max 119.28 / n=4

gemma-4-31b-it

31B / Q4_K_M

decode failed

n=4, throughput 0 tok/s

Decode failed: model prefilled then produced 0 tokens across all runs on AMD ROCm 7.10 + Ollama 0.24.0. A real, reproducible failure mode, not missing data.

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $999 as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.