Consumer GPUNVIDIA

NVIDIA GeForce RTX 5080 16GB

Curated Aggregate·2026-02-0811 workloads · 15 records
MyAI Rating9.1emb/s · Embedding throughput

VRAM

16 GB

TDP

360 W

MSRP

$999

Perf/W

0.46 emb/s/W

Cost/1K tok

$0.06/M

Tested

2026-02-08

Quick answer

How many tokens per second does NVIDIA GeForce RTX 5080 16GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 5080 16GB produces approximately 7.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 16GB VRAM, 360W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-rtx5080-l3-70b-q4)As of 2026-02-15

Overview

Enthusiast Tier

NVIDIA GeForce RTX 5080 is the Blackwell-architecture enthusiast card released in 2025. 16GB GDDR7 at ~960 GB/s bandwidth, 360W TDP. Blackwell compute with FP4 support but capped at 16GB VRAM — the entry point for Blackwell consumer AI.

AI Usefulness

16GB VRAM limits full-GPU models to ~13B-class at Q4. 32B models require aggressive quantization or partial offload. Real-world: Treydol-LLM-Asure-12B Q4 at ~70 tok/s (verified-lab on this device). Best for 7-13B daily use, coding agents on smaller models, and as a secondary GPU in multi-card rigs. Overkill for sub-13B if you're budget-conscious.

Editorial Verdict

7.5
Editor Rating

Buy for mixed gaming+AI or as a secondary GPU in multi-card rigs. Skip if your primary use case is local AI — used 3090 gives you 24GB at lower cost, and 5070 Ti 16GB gives similar VRAM for $250 less.

What it does well

  • +16GB GDDR7 at ~960 GB/s — entry point for Blackwell AI
  • +FP4 support for future quantized models
  • +360W TDP — reasonable power draw for the performance tier
  • +Measured 69.9 tok/s (σ=1.46) on Trendyol-LLM-Asure 12B Q4 — verified-lab data

Where it breaks

  • , 16GB VRAM caps at 13B-class full-GPU — the same ceiling as 4080 Super
  • , 32B models require aggressive quantization or partial offload
  • , Only 16GB VRAM at $999 — used 3090 (24GB) at $600-800 is better $/VRAM
  • , Overkill for sub-13B if you're budget-conscious

Sweet Spot

13B-class Q4 full-GPU at 70+ tok/s — excellent for coding agents on Qwen 2.5 Coder 14B or similar

Bad Use Cases

  • ×32B-class as daily driver (need 24GB+ card)
  • ×70B inference without second GPU
  • ×Pure VRAM capacity builds (used 3090 wins on $/GB)
  • ×Budget-constrained AI builds (4060 Ti 16GB is half the price)

What Breaks First

VRAM ceiling — 16GB fills fast with 32B models at useful context. KV cache eats the headroom quickly.

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2LM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Ollama + Qwen 2.5 Coder 14B Q4_K_M for coding
  • llama.cpp + Gemma 4 9B Q4 (~131 tok/s measured on this device)
  • 750W Gold PSU + standard ATX case

Power & Cooling

360W TGP, ~250-300W sustained decode. 750W PSU minimum. Reasonable thermals in standard cases. The most power-efficient Blackwell card for AI.

Verdict

NVIDIA GeForce RTX 5080 16GB with 16GB VRAM at 360W TDP, scored across 11 workloads with 15 benchmark records.

Best workload

Embedding throughput

8200 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

04590135180Llama 38B Q4Llama 38B FP16Llama 370B Q4DeepSeek-R17BQwen 2.514BMistral 7B

Benchmarks (11 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

165.0tok/sQ4_K_M4K±4.8Curated Aggregate2026-02-08
Llama 3 8B FP16

llm

88.0tok/sFP164K, Curated Aggregate2026-02-12
Llama 3 70B Q4

llm

7.0tok/sQ4_K_M4K, Curated Aggregate2026-02-15
DeepSeek-R1 7B

llm

138.0tok/sQ4_K_M8K, Curated Aggregate2026-02-20
Qwen 2.5 14B

llm

86.0tok/sQ4_K_M8K, Curated Aggregate2025-03-16
SDXL image gen

image

18.0img/minFP16, , Curated Aggregate2026-02-25
Whisper transcription

audio

95.0x RTFP16, , Curated Aggregate2026-03-01
Embedding throughput

embedding

8200.0emb/sFP161K, Curated Aggregate2025-03-04
Mistral 7B

llm

122.0tok/sQ4_K_M4K, Curated Aggregate2025-02-15
Gemma 2 9B

llm

105.0tok/sQ4_K_M8K, Curated Aggregate2025-02-25
Asure 12B

llm

69.9tok/sQ4_K_M8K±1.5Lab Verified2025-05-01

MyAI Score

Reference
9.1/10

NVIDIA GeForce RTX 5080 16GB clears a 9.1/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
292
Capability
129
Efficiency
50
Value
78
Trust
54
Coverage
60
Composite benchmark663 / 1000

Workload Fit

What models fit this 16GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Tight
13-14B
Excellent
Good
Won't fit
32B
Tight
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Top Benchmarks
Embedding throughput8200 emb/s
Llama 3 8B Q4165 tok/s
DeepSeek-R1 7B138 tok/s

Source

TEI on Blackwell.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.1

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2025-03-04; 541 days old.

Open primary source

Verified Local-Lab Run

Measured on your own machine via ollama-local-api on 2026-05-28.

Batch=1, 2048 context, 512 generated tokens, consecutive warm-model runs. Generation TPS uses Ollama eval_count/eval_duration.

kumru-2b

2.4B / Q4_K_M

429.62 tok/s

median 443.66 / min 397.42 / max 454.33 / n=5

brooqs-mistral-turkish-v2-latest

7.2B / Q4_0

160.83 tok/s

median 161.11 / min 159.88 / max 161.79 / n=5

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

Search Amazon listings
  • +Use $999 as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.