Consumer GPUNVIDIA

NVIDIA GeForce RTX 3090 24GB

Curated Aggregate·2025-02-2510 workloads · 15 records
MyAI Rating8.6emb/s · Embedding throughput

VRAM

24 GB

TDP

350 W

MSRP

$1.5k

Perf/W

0.25 emb/s/W

Cost/1K tok

$0.18/M

Tested

2025-02-25

Quick answer

How many tokens per second does NVIDIA GeForce RTX 3090 24GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 3090 24GB produces approximately 9.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 24GB VRAM, 350W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-rtx3090-l3-70b-q4)As of 2025-04-25

Overview

Value VRAM King

NVIDIA GeForce RTX 3090 24GB — Ampere flagship from 2020. 24GB GDDR6X at ~936 GB/s, 350W TDP. Same VRAM capacity as RTX 4090 but ~60% the throughput on AI workloads. No FP8 acceleration.

AI Usefulness

The value king for VRAM-hungry local AI. 24GB runs 32B Q4 full-GPU at ~35-45 tok/s. Used market pricing (~$600-800) makes it the best $/VRAM in consumer hardware. NVLink support on some models enables dual-3090 rigs with 48GB combined VRAM — the budget workstation builder's secret weapon. Slower on prompt processing than Ada/Blackwell but decode speed is bandwidth-bound and competitive.

Editorial Verdict

8.5
Editor Rating

Buy used if you can find a clean unit under $700. The value-for-VRAM proposition is unmatched — nothing else gives you 24GB at this price. Pair with a second used 3090 for 48GB combined VRAM at ~$1,400 total.

What it does well

  • +24GB GDDR6X at ~936 GB/s — same VRAM as RTX 4090 at ~60% the throughput
  • +Used market pricing ($600-800) — best $/VRAM in consumer hardware
  • +NVLink support on some models enables dual-3090 rigs with 48GB combined
  • +Mature Ampere architecture with universal CUDA support

Where it breaks

  • , ~60% the tok/s of 4090 on same model — Ampere is slower on prompt processing
  • , 350W TDP — older architecture, less efficient than Ada/Blackwell
  • , No FP8 acceleration — missing modern quantization features
  • , Used market risks — mining cards, worn fans, no warranty

Sweet Spot

32B Q4 full-GPU at 35-45 tok/s — or dual 3090s for 70B Q4 at 25-35 tok/s with 48GB combined

Bad Use Cases

  • ×New-card buyers who want warranty
  • ×Maximum tok/s on small models (Ada/Blackwell cards are faster)
  • ×Single-slot builds (2.5-3 slot card)
  • ×Compact ITX cases without serious airflow

What Breaks First

VRAM thermals under sustained load — GDDR6X runs hot. Repadding recommended for used cards. Fan bearings on mining cards typically the first mechanical failure.

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2LM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Dual 3090 via NVLink for 48GB combined — layer-split 70B Q4 at 25-35 tok/s
  • Ollama + Qwen 2.5 Coder 32B Q4 for budget coding agent
  • 850W+ PSU per card, 1200W+ for dual setup

Power & Cooling

350W TGP. Used cards may draw more due to degraded thermal paste. Undervolt recommended for 24/7 operation. ~$10/month electricity for single card at 4hrs/day US average.

Verdict

NVIDIA GeForce RTX 3090 24GB with 24GB VRAM at 350W TDP, scored across 10 workloads with 15 benchmark records.

Best workload

Embedding throughput

4900 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0255075100Llama 38B Q4Llama 370B Q4Qwen 2.514BGemma 29BDeepSeek-R17BMistral 7B

Benchmarks (10 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

88.0tok/sQ4_K_M4K, Curated Aggregate2025-02-25
Llama 3 70B Q4

llm

9.0tok/sQ4_K_M4K, Curated Aggregate2025-04-25
SDXL image gen

image

12.0img/minFP16, , Curated Aggregate2024-12-10
Whisper transcription

audio

56.0x RTFP16, , Curated Aggregate2024-08-17
Qwen 2.5 14B

llm

55.0tok/sQ4_K_M8K, Curated Aggregate2025-03-08
Gemma 2 9B

llm

65.0tok/sQ4_K_M8K, Curated Aggregate2024-06-12
DeepSeek-R1 7B

llm

82.0tok/sQ4_K_M8K, Curated Aggregate2025-01-29
Embedding throughput

embedding

4900.0emb/sFP161K, Curated Aggregate2024-05-18
Mistral 7B

llm

78.0tok/sQ4_K_M4K, Curated Aggregate2024-04-22
Asure 12B

llm

38.0tok/sQ4_K_M8K±2.3Curated Aggregate2025-05-01

MyAI Score

Flagship
8.6/10

NVIDIA GeForce RTX 3090 24GB clears a 8.6/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
268
Capability
140
Efficiency
43
Value
49
Trust
44
Coverage
60
Composite benchmark604 / 1000

Workload Fit

What models fit this 24GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Tight
32B
Excellent
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Top Benchmarks
Embedding throughput4900 emb/s
Llama 3 8B Q488 tok/s
DeepSeek-R1 7B82 tok/s

Public Trust Layer

Trust score

6/10

MyAI rating

8.6

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-05-18; 831 days old.

Open primary source

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

Search Amazon listings
  • +Use $1.5k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

This is a strong buy when the listing stays close to reference pricing and matches the workload you actually run.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.