Consumer GPUNVIDIA

NVIDIA GeForce RTX 4090 24GB

Curated Aggregate·2025-10-0212 workloads · 23 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

24 GB

TDP

450 W

MSRP

$1.6k

Perf/W

0.29 tok/s/W

Cost/1K tok

$0.13/M

Tested

2025-10-02

Quick answer

How many tokens per second does NVIDIA GeForce RTX 4090 24GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 4090 24GB has a source-attributed result of 14.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-rtx4090-l3-70b-q4)As of 2026-03-08

Overview

Workstation Default

NVIDIA GeForce RTX 4090 is the Ada Lovelace flagship released in 2022. 24GB GDDR6X at ~1 TB/s bandwidth, 450W TDP, 4th-gen Tensor Cores with FP8 acceleration. The community-default high-end local-AI card from 2022-2025 with universal CUDA software support.

AI Usefulness

24GB suits many 8B–32B quantized models, subject to artifact size and context. 13B FP16 weights are about 26GB; 70B Q3 and Q4 also exceed this card before runtime overhead. Larger models require a supported split across devices or CPU offload; no speed is promised for an unspecified configuration.

Editorial Verdict

9.4
Editor Rating

Buy if you want the best single-card experience for 32B-class models and either bought at MSRP or found a clean used unit. Skip if 70B is your daily target or if 16GB cards cover your needs.

What it does well

  • +24GB GDDR6X at 1 TB/s — the community-default high-end local-AI card
  • +Universal CUDA support, mature drivers, every engine has a happy path
  • +Used market pricing ($1,400-1,700) makes it compelling vs 5090 MSRP
  • +450W TDP manageable with quality 850W PSU

Where it breaks

  • , 24GB VRAM caps at 32B-class full-GPU — 70B requires partial offload at 22-28 tok/s
  • , 2022 silicon — no FP4, no Blackwell features
  • , Warranty terms vary by board vendor; check the actual warranty
  • , Resale premium persists even after 5090 launch

Sweet Spot

32B-class Q4 full-GPU at 70+ tok/s — Qwen 3 32B, Qwen 2.5 Coder 32B, QwQ 32B all fit comfortably

Bad Use Cases

  • ×Full-residency 70B Q4 on one 24GB card
  • ×Genuine 100B+ MoE workloads
  • ×Maximum tok/s on sub-7B models (lower-end cards are better $/throughput)
  • ×Production 24/7 serving (consumer warranty, no ECC)

What Breaks First

VRAM at 32K+ context — 70B Q4 at 32K context uses ~20GB for KV cache alone. Budget context length carefully.

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2TensorRT-LLMLM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Ollama + Qwen 2.5 Coder 32B Q4_K_M for solo coding agent
  • ExLlamaV2 + EXL2 4.65bpw 32B for single-stream throughput king
  • Ubuntu 24.04 + CUDA 12.4 + Open WebUI Docker for homelab default
  • 1000W Gold PSU + three-fan tower or AIO cooling

Power & Cooling

450W TGP, ~320-380W sustained decode. 850W PSU minimum, 1000W recommended. ~$8-10/month electricity at US average for 4hrs/day. Undervolt -100mV for zero perf cost thermal improvement.

Verdict

NVIDIA GeForce RTX 4090 24GB with 24GB VRAM at 450W TDP, scored across 12 workloads with 23 benchmark records.

Reference workload

DeepSeek-R1 7B

125 tok/s

Quantization

Q4_K_M

8K context · batch 1

LLM Inference Performance

04080120160Llama 3 8BQ4Llama 3 8BFP16Llama 370B Q4Mistral 7BDeepSeek-R17BGemma 29B

Benchmarks (12 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

132.0tok/sQ4_K_M4K, Curated Aggregate2025-10-02
Llama 3 8B FP16

llm

72.0tok/sFP164K, Curated Aggregate2024-12-22
Llama 3 70B Q4

llm

14.0tok/sQ4_K_M4K, Curated Aggregate2026-03-08
Mistral 7B

llm

145.0tok/sQ4_K_M4K, Curated Aggregate2026-03-19
SDXL image gen

image

24.0img/minFP16, , Curated Aggregate2024-06-22
Whisper transcription

audio

78.0x RTFP16, , Curated Aggregate2025-03-08
Embedding throughput

embedding

5800.0emb/sFP161K, Curated Aggregate2025-09-04
DeepSeek-R1 7B

llm

125.0tok/sQ4_K_M8K, Curated Aggregate2026-05-25
Gemma 2 9B

llm

92.0tok/sQ4_K_M8K, Curated Aggregate2025-08-25
Qwen 2.5 14B

llm

68.0tok/sQ4_K_M8K, Curated Aggregate2025-07-04
Phi-3 Mini

llm

285.0tok/sQ4_K_M4K, Curated Aggregate2025-08-12
Asure 12B

llm

98.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA GeForce RTX 4090 24GB. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 24GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Won't fit
32B
Good
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Reported batch-one examples
DeepSeek-R1 7B125 tok/s
Mistral 7B145 tok/s
Llama 3 70B Q414 tok/s

Source

llama.cpp b3850, refreshed seed.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Aging

Source-linked row with explicit verification status.

Record date: 2026-05-25; 107 days old.

Open primary source

Best place to start

Where to buy

Most buyers

Retailer we'd check first

Amazon

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

View current Amazon listing
  • +Use $1.6k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: this button opens the mapped Amazon product listing for this device. We may earn from qualifying purchases.