Consumer GPUNVIDIA

NVIDIA GeForce RTX 5090 32GB

Curated Aggregate·2025-12-3030 workloads · 47 records
MyAI Rating9.7emb/s · Embedding throughput

VRAM

32 GB

TDP

575 W

MSRP

$2.0k

Perf/W

0.32 emb/s/W

Cost/1K tok

$0.12/M

Tested

2025-12-30

Quick answer

How many tokens per second does NVIDIA GeForce RTX 5090 32GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 5090 32GB produces approximately 28.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 32GB VRAM, 575W TDP). That figure comes from 5 measured runs on llama.cpp b4876. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-rtx5090-l3-70b-q4)As of 2025-10-10

Overview

Consumer Flagship

NVIDIA GeForce RTX 5090 is the Blackwell-architecture consumer flagship GPU released in 2025. It features 32GB of GDDR7 VRAM on a 512-bit bus delivering ~1.79 TB/s memory bandwidth, 575W TDP, and native FP4 acceleration via 5th-gen Tensor Cores. Built on TSMC 4NP process with 92 billion transistors.

AI Usefulness

The 32GB VRAM is the first consumer card to comfortably host 70B-class models at Q4 with usable context headroom. Expect ~40-55 tok/s on Llama 3.3 70B Q4 full-GPU. Ideal for single-card 70B inference, 32B FP16 with full context, and concurrent image-gen + LLM workloads. FP4 acceleration is emerging but engine support still maturing through 2026.

Editorial Verdict

9.6
Editor Rating

Buy if 70B Q4 full-GPU is your daily target and you can find one near MSRP. Skip if 4090 covers your model range or if the scalper premium exceeds 20%.

What it does well

  • +32GB GDDR7 at 1.79 TB/s — first consumer card to comfortably host 70B Q4 full-GPU
  • +Blackwell FP4 acceleration future-proofs for next-gen quantized models
  • +Roughly 1.8× the bandwidth of RTX 4090, decode speed scales linearly
  • +Universal CUDA support across every inference engine

Where it breaks

  • , 575W TDP demands 1000W+ PSU — not a drop-in upgrade
  • , Still scalper-priced at $2,300-2,800 through mid-2026
  • , Overkill for 32B-class models — 4090 does the same job for less
  • , No NVLink — multi-GPU goes over PCIe

Sweet Spot

70B Q4 full-GPU at 40-55 tok/s with 8-16K context — the first consumer card where 70B daily-driving is practical

Bad Use Cases

  • ×32B-class only workloads (4090 is better $/token)
  • ×Sub-1000W PSU builds
  • ×Compact ITX cases without aggressive cooling
  • ×Anyone price-sensitive enough that scalper premium hurts

What Breaks First

12V-2x6 connector under sustained 575W load — use native cable with 35mm straight run before any bend

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2TensorRT-LLMLM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • vLLM + 70B AWQ-INT4 for multi-user serving
  • ExLlamaV2 + EXL2 4.65bpw 70B for single-stream king
  • Ubuntu 24.04 + CUDA 12.6+ + Open WebUI Docker
  • 1200W Platinum PSU + mid-tower with direct GPU airflow

Power & Cooling

575W TGP sustained under inference. Budget 1200W PSU minimum. Consider undervolting for 24/7 operation. Adds ~$15-20/month to electricity at US average rates.

Verdict

NVIDIA GeForce RTX 5090 32GB with 32GB VRAM at 575W TDP, scored across 30 workloads with 47 benchmark records.

Best workload

Embedding throughput

12200 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0200400600800Llama 38B Q4Llama 38B FP16Llama 370B Q4Qwen 2.514BDeepSeek-R17BGemma 29B

Benchmarks (30 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

720.0tok/sQ4_K_M4K, Curated Aggregate2025-03-12
Llama 3 8B FP16

llm

96.0tok/sFP164K, Curated Aggregate2024-09-04
Llama 3 70B Q4

llm

78.0tok/sQ4_K_M4K, Curated Aggregate2025-03-25
Qwen 2.5 14B

llm

115.0tok/sQ4_K_M8K, Curated Aggregate2025-03-02
DeepSeek-R1 7B

llm

185.0tok/sQ4_K_M8K, Curated Aggregate2025-02-22
SDXL image gen

image

38.0img/minFP16, , Curated Aggregate2024-06-29
Gemma 2 9B

llm

152.0tok/sQ4_K_M8K, Curated Aggregate2025-02-18
Whisper transcription

audio

128.0x RTFP16, , Curated Aggregate2026-03-12
Embedding throughput

embedding

12200.0emb/sFP161K, Curated Aggregate2025-04-02
Mistral 7B

llm

198.0tok/sQ4_K_M4K, Curated Aggregate2026-02-12
Phi-3 Mini

llm

348.0tok/sQ4_K_M4K, Curated Aggregate2026-02-22
Llama 3 70B Q8

llm

14.0tok/sQ8_04K, Curated Aggregate2026-04-08
Asure 12B

llm

85.0tok/sQ4_K_M8K±4.2Curated Aggregate2025-05-01
Gemma 4 31B Dense

llm

58.8tok/sQ4_K_M4K, Community Verified2026-07-20
Muse Glimmer 30B

llm

80.1tok/sQ4_K_M4K, Community Verified2026-08-10
Gemma 4 26B-A4B

llm

148.9tok/sQ4_K_M4K, Community Verified2026-07-20
North Mini Code 1.0

llm

177.0tok/sQ4_K_M4K, Community Verified2026-07-21
Ornith 1.0 35B

llm

174.9tok/sQ4_K_M4K, Community Verified2026-07-21
Qwen3 Coder 30B-A3B

llm

229.7tok/sQ4_K_M4K, Community Verified2026-07-20
Qwen3.5 35B-A3B

llm

173.5tok/sQ4_K_M4K, Community Verified2026-07-21
Laguna XS 2.1

llm

235.1tok/sQ4_K_M4K, Community Verified2026-07-21
Qwen3.5 27B

llm

67.9tok/sQ4_K_M4K, Community Verified2026-07-20
Qwen3.6 27B

llm

67.9tok/sQ4_K_M4K, Community Verified2026-07-20
Granite 4.1 30B

llm

73.3tok/sQ4_K_M4K, Community Verified2026-07-20
Nemotron 3 Nano Omni 33B

llm

248.6tok/sQ4_K_M4K, Community Verified2026-07-21
GLM-4.7-Flash

llm

217.5tok/sQ4_K_M4K, Community Verified2026-07-21
Qwen3.6 35B-A3B

llm

169.9tok/sQ4_K_M4K, Community Verified2026-07-21
Mellum2 12B-A2.5B

llm

419.6tok/sQ4_K_M4K, Community Verified2026-07-20
Dolphin 3.0 8B

llm

197.8tok/sQ4_K_M4K, Community Verified2026-07-20
Hermes 3 Llama 3.1 8B

llm

204.8tok/sQ4_K_M4K, Community Verified2026-07-20

MyAI Score

Mythic
9.7/10

NVIDIA GeForce RTX 5090 32GB clears a 9.7/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
378
Capability
165
Efficiency
84
Value
92
Trust
66
Coverage
60
Composite benchmark845 / 1000

Workload Fit

What models fit this 32GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Good
32B
Excellent
Tight
Won't fit
70B
Tight
Won't fit
Won't fit
Top Benchmarks
Embedding throughput12200 emb/s
Llama 3 8B Q4720 tok/s
Mellum2 12B-A2.5B420 tok/s

Source

TEI 1.5, batch 32.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.7

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2025-04-02; 512 days old.

Open primary source

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $2.0k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.