Consumer GPUNVIDIA

NVIDIA GeForce RTX 5090 32GB

Curated Aggregate·2025-12-3030 workloads · 47 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

32 GB

TDP

575 W

MSRP

$2.0k

Perf/W

0.32 tok/s/W

Cost/1K tok

$0.12/M

Tested

2025-12-30

Quick answer

How many tokens per second does NVIDIA GeForce RTX 5090 32GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 5090 32GB has a source-attributed result of 28.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-rtx5090-l3-70b-q4)As of 2025-10-10

Overview

Consumer Flagship

NVIDIA GeForce RTX 5090 is the Blackwell-architecture consumer flagship GPU released in 2025. It features 32GB of GDDR7 VRAM on a 512-bit bus delivering ~1.79 TB/s memory bandwidth, 575W TDP, and native FP4 acceleration via 5th-gen Tensor Cores. Built on TSMC 4NP process with 92 billion transistors.

AI Usefulness

32GB suits many 8B–32B quantized models. Llama 3 70B Q4_K_M weights alone are roughly 40GB, and 32B FP16 weights roughly 64GB; neither fits fully in 32GB. Use a smaller artifact, multiple GPUs or explicit CPU offload, allowing for KV cache and runtime memory. Compare only documented model and runtime settings.

Editorial Verdict

9.6
Editor Rating

Consider it for a 32GB CUDA workload after checking the artifact and total system cost. Choose more GPU memory for full-residency 70B Q4.

What it does well

  • +32GB GDDR7 at approximately 1.79TB/s
  • +CUDA support with runtime-specific Blackwell requirements

Where it breaks

  • , 575W TDP demands 1000W+ PSU — not a drop-in upgrade
  • , Still scalper-priced at $2,300-2,800 through mid-2026
  • , Overkill for 32B-class models — 4090 does the same job for less
  • , No NVLink — multi-GPU goes over PCIe

Sweet Spot

8B–32B quantized models with room reserved for runtime and KV cache.

Bad Use Cases

  • ×32B-class only workloads (4090 is better $/token)
  • ×Sub-1000W PSU builds
  • ×Compact ITX cases without aggressive cooling
  • ×Anyone price-sensitive enough that scalper premium hurts

What Breaks First

12V-2x6 connector under sustained 575W load — use native cable with 35mm straight run before any bend

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2TensorRT-LLMLM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Qwen 2.5 Coder 32B Q4 with a supported CUDA runtime and measured memory use
  • A PSU, case and native GPU cable approved for the exact board

Power & Cooling

575W is the GPU power rating, not a measured constant inference draw. Electricity cost = measured system watts / 1000 × hours × your tariff.

Verdict

NVIDIA GeForce RTX 5090 32GB with 32GB VRAM at 575W TDP, scored across 30 workloads with 47 benchmark records.

Reference workload

Muse Glimmer 30B

80.1 tok/s

Quantization

Q4_K_M

4K context · batch 1

LLM Inference Performance

050100150200Llama 3 8BQ4Llama 3 8BFP16Llama 370B Q4Qwen 2.514BDeepSeek-R17BGemma 29B

Benchmarks (30 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

182.0tok/sQ4_K_M4K, Curated Aggregate2025-12-30
Llama 3 8B FP16

llm

96.0tok/sFP164K, Curated Aggregate2024-09-04
Llama 3 70B Q4

llm

28.0tok/sQ4_K_M4K, Curated Aggregate2025-10-10
Qwen 2.5 14B

llm

105.0tok/sQ4_K_M8K, Curated Aggregate2026-05-24
DeepSeek-R1 7B

llm

185.0tok/sQ4_K_M8K, Curated Aggregate2025-02-22
SDXL image gen

image

32.0img/minFP16, , Curated Aggregate2026-03-25
Gemma 2 9B

llm

122.0tok/sQ4_K_M8K, Curated Aggregate2026-03-05
Whisper transcription

audio

128.0x RTFP16, , Curated Aggregate2026-03-12
Embedding throughput

embedding

9100.0emb/sFP161K, Curated Aggregate2026-05-22
Mistral 7B

llm

182.0tok/sQ4_K_M8K, Curated Aggregate2026-02-15
Phi-3 Mini

llm

320.0tok/sQ4_K_M4K, Curated Aggregate2026-03-01
Llama 3 70B Q8

llm

14.0tok/sQ8_04K, Curated Aggregate2026-04-08
Asure 12B

llm

140.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01
Gemma 4 31B Dense

llm

58.8tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Muse Glimmer 30B

llm

80.1tok/sQ4_K_M4K, Curated Aggregate2026-08-10
Gemma 4 26B-A4B

llm

148.9tok/sQ4_K_M4K, Curated Aggregate2026-07-20
North Mini Code 1.0

llm

177.0tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Ornith 1.0 35B

llm

174.9tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Qwen3 Coder 30B-A3B

llm

229.7tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Qwen3.5 35B-A3B

llm

173.5tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Laguna XS 2.1

llm

235.1tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Qwen3.5 27B

llm

67.9tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Qwen3.6 27B

llm

67.9tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Granite 4.1 30B

llm

73.3tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Nemotron 3 Nano Omni 33B

llm

248.6tok/sQ4_K_M4K, Curated Aggregate2026-07-21
GLM-4.7-Flash

llm

217.5tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Qwen3.6 35B-A3B

llm

169.9tok/sQ4_K_M4K, Curated Aggregate2026-07-21
Mellum2 12B-A2.5B

llm

419.6tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Dolphin 3.0 8B

llm

197.8tok/sQ4_K_M4K, Curated Aggregate2026-07-20
Hermes 3 Llama 3.1 8B

llm

204.8tok/sQ4_K_M4K, Curated Aggregate2026-07-20

Test Environment

Framework

llama.cpp 4dee52f

Runs

5

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA GeForce RTX 5090 32GB. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 32GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Tight
32B
Excellent
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Reported batch-one examples
Muse Glimmer 30B80 tok/s
North Mini Code 1.0177 tok/s
Ornith 1.0 35B175 tok/s

Source

Owner-measured on a rented RTX 5090 (Vast.ai) under llama.cpp @4dee52f, Q4_K_M (16.7GB VRAM). Meta's open agentic model (Apache-2.0), day-one. HumanEval+ 91.5% / base 97.6% (EvalPlus, 164 tasks) — rank 3 of 16, best non-code-tuned. DFlash spec-decode gave no measurable speedup.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

5

Freshness

Recent

Source-linked row with explicit verification status.

Record date: 2026-08-10; 30 days old.

Open primary source

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $2.0k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.