Consumer GPUNVIDIA

NVIDIA GeForce RTX 4090 24GB

Curated Aggregate·2025-10-0212 workloads · 23 records
MyAI Rating9.1emb/s · Embedding throughput

VRAM

24 GB

TDP

450 W

MSRP

$1.6k

Perf/W

0.29 emb/s/W

Cost/1K tok

$0.13/M

Tested

2025-10-02

Quick answer

How many tokens per second does NVIDIA GeForce RTX 4090 24GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 4090 24GB produces approximately 14.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 24GB VRAM, 450W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-rtx4090-l3-70b-q4)As of 2026-03-08

Overview

Workstation Default

NVIDIA GeForce RTX 4090 is the Ada Lovelace flagship released in 2022. 24GB GDDR6X at ~1 TB/s bandwidth, 450W TDP, 4th-gen Tensor Cores with FP8 acceleration. The community-default high-end local-AI card from 2022-2025 with universal CUDA software support.

AI Usefulness

24GB VRAM comfortably runs 32B-class models at Q4 (Qwen 3 32B, Qwen 2.5 Coder 32B) at 70+ tok/s. 70B Q4 requires partial offload to system RAM, dropping to ~22-28 tok/s. The sweet-spot card for solo coding agents and 32B-class daily drivers. Used market pricing makes it compelling vs MSRP 5090.

Editorial Verdict

9.4
Editor Rating

Buy if you want the best single-card experience for 32B-class models and either bought at MSRP or found a clean used unit. Skip if 70B is your daily target or if 16GB cards cover your needs.

What it does well

  • +24GB GDDR6X at 1 TB/s — the community-default high-end local-AI card
  • +Universal CUDA support, mature drivers, every engine has a happy path
  • +Used market pricing ($1,400-1,700) makes it compelling vs 5090 MSRP
  • +450W TDP manageable with quality 850W PSU

Where it breaks

  • , 24GB VRAM caps at 32B-class full-GPU — 70B requires partial offload at 22-28 tok/s
  • , 2022 silicon — no FP4, no Blackwell features
  • , Consumer warranty terms — 24/7 inference technically out of spec
  • , Resale premium persists even after 5090 launch

Sweet Spot

32B-class Q4 full-GPU at 70+ tok/s — Qwen 3 32B, Qwen 2.5 Coder 32B, QwQ 32B all fit comfortably

Bad Use Cases

  • ×70B daily-driving (5090 or dual-3090 are better fits)
  • ×Genuine 100B+ MoE workloads
  • ×Maximum tok/s on sub-7B models (lower-end cards are better $/throughput)
  • ×Production 24/7 serving (consumer warranty, no ECC)

What Breaks First

VRAM at 32K+ context — 70B Q4 at 32K context uses ~20GB for KV cache alone. Budget context length carefully.

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2TensorRT-LLMLM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Ollama + Qwen 2.5 Coder 32B Q4_K_M for solo coding agent
  • ExLlamaV2 + EXL2 4.65bpw 32B for single-stream throughput king
  • Ubuntu 24.04 + CUDA 12.4 + Open WebUI Docker for homelab default
  • 1000W Gold PSU + three-fan tower or AIO cooling

Power & Cooling

450W TGP, ~320-380W sustained decode. 850W PSU minimum, 1000W recommended. ~$8-10/month electricity at US average for 4hrs/day. Undervolt -100mV for zero perf cost thermal improvement.

Verdict

NVIDIA GeForce RTX 4090 24GB with 24GB VRAM at 450W TDP, scored across 12 workloads with 23 benchmark records.

Best workload

Embedding throughput

8500 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0150300450600Llama 38B Q4Llama 38B FP16Llama 370B Q4Mistral 7BDeepSeek-R17BGemma 29B

Benchmarks (12 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

540.0tok/sQ4_K_M4K, Curated Aggregate2024-08-21
Llama 3 8B FP16

llm

72.0tok/sFP164K, Curated Aggregate2024-12-22
Llama 3 70B Q4

llm

14.0tok/sQ4_K_M4K, Curated Aggregate2026-03-08
Mistral 7B

llm

145.0tok/sQ4_K_M4K, Curated Aggregate2026-03-19
SDXL image gen

image

24.0img/minFP16, , Curated Aggregate2024-06-22
Whisper transcription

audio

78.0x RTFP16, , Curated Aggregate2025-03-08
Embedding throughput

embedding

8500.0emb/sFP161K, Curated Aggregate2024-07-15
DeepSeek-R1 7B

llm

125.0tok/sQ4_K_M8K, Curated Aggregate2026-05-25
Gemma 2 9B

llm

108.0tok/sQ4_K_M8K, Curated Aggregate2024-11-02
Qwen 2.5 14B

llm

78.0tok/sQ4_K_M8K, Curated Aggregate2024-11-19
Phi-3 Mini

llm

285.0tok/sQ4_K_M4K, Curated Aggregate2025-08-12
Asure 12B

llm

60.0tok/sQ4_K_M8K±3.2Curated Aggregate2025-05-01

MyAI Score

Reference
9.1/10

NVIDIA GeForce RTX 4090 24GB clears a 9.1/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
303
Capability
157
Efficiency
47
Value
58
Trust
48
Coverage
60
Composite benchmark673 / 1000

Workload Fit

What models fit this 24GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Tight
32B
Excellent
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Top Benchmarks
Embedding throughput8500 emb/s
Llama 3 8B Q4540 tok/s
Phi-3 Mini285 tok/s

Source

TEI 1.5 batch 32.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.1

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-07-15; 773 days old.

Open primary source

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $1.6k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.