Consumer GPUNVIDIA

NVIDIA GeForce RTX 5080 16GB

Curated Aggregate·2026-02-0811 workloads · 15 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

16 GB

TDP

360 W

MSRP

$999

Perf/W

0.46 emb/s/W

Cost/1K tok

$0.06/M

Tested

2026-02-08

Quick answer

How many tokens per second does NVIDIA GeForce RTX 5080 16GB produce on Llama 3 70B Q4?

NVIDIA GeForce RTX 5080 16GB has a source-attributed result of 7.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-rtx5080-l3-70b-q4)As of 2026-02-15

Overview

Enthusiast Tier

NVIDIA GeForce RTX 5080 is the Blackwell-architecture enthusiast card released in 2025. 16GB GDDR7 at ~960 GB/s bandwidth, 360W TDP. Blackwell compute with FP4 support but capped at 16GB VRAM — the entry point for Blackwell consumer AI.

AI Usefulness

16GB suits many 7B–14B Q4 models. A typical 32B Q4 artifact needs more memory, before context and runtime overhead. The Asure throughput entry is a reported result with no public run artifact, not independently verified lab evidence.

Editorial Verdict

7.5
Editor Rating

Buy for mixed gaming+AI or as a secondary GPU in multi-card rigs. Skip if your primary use case is local AI — used 3090 gives you 24GB at lower cost, and 5070 Ti 16GB gives similar VRAM for $250 less.

What it does well

  • +16GB GDDR7
  • +Blackwell compute with runtime-specific support
  • +360W GPU power rating

Where it breaks

  • , 16GB VRAM caps at 13B-class full-GPU — the same ceiling as 4080 Super
  • , 32B models require aggressive quantization or partial offload
  • , Only 16GB VRAM at $999 — used 3090 (24GB) at $600-800 is better $/VRAM
  • , Overkill for sub-13B if you're budget-conscious

Sweet Spot

13B-class Q4 full-GPU at 70+ tok/s — excellent for coding agents on Qwen 2.5 Coder 14B or similar

Bad Use Cases

  • ×32B-class as daily driver (need 24GB+ card)
  • ×70B inference without second GPU
  • ×Pure VRAM capacity builds (used 3090 wins on $/GB)
  • ×Budget-constrained AI builds (4060 Ti 16GB is half the price)

What Breaks First

VRAM ceiling — 16GB fills fast with 32B models at useful context. KV cache eats the headroom quickly.

Software Support

Ollamallama.cppvLLMSGLangExLlamaV2LM StudioPyTorch

Ubuntu 24.04 (excellent), Windows 11 (excellent), WSL2 (excellent), macOS (unsupported)

Best Pairings

  • Qwen 2.5 Coder 14B Q4 with a supported runtime
  • A compatible PSU and case selected for the exact board

Power & Cooling

360W TGP, ~250-300W sustained decode. 750W PSU minimum. Reasonable thermals in standard cases. The most power-efficient Blackwell card for AI.

Verdict

NVIDIA GeForce RTX 5080 16GB with 16GB VRAM at 360W TDP, scored across 11 workloads with 15 benchmark records.

Reference workload

Embedding throughput

7200 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

04590135180Llama 3 8BQ4Llama 3 8BFP16Llama 370B Q4DeepSeek-R17BQwen 2.514BMistral 7B

Benchmarks (11 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

165.0tok/sQ4_K_M4K, Curated Aggregate2026-02-08
Llama 3 8B FP16

llm

88.0tok/sFP164K, Curated Aggregate2026-02-12
Llama 3 70B Q4

llm

7.0tok/sQ4_K_M4K, Curated Aggregate2026-02-15
DeepSeek-R1 7B

llm

138.0tok/sQ4_K_M8K, Curated Aggregate2026-02-20
Qwen 2.5 14B

llm

82.0tok/sQ4_K_M8K, Curated Aggregate2026-02-22
SDXL image gen

image

18.0img/minFP16, , Curated Aggregate2026-02-25
Whisper transcription

audio

95.0x RTFP16, , Curated Aggregate2026-03-01
Embedding throughput

embedding

7200.0emb/sFP161K, Curated Aggregate2026-05-23
Mistral 7B

llm

122.0tok/sQ4_K_M4K, Curated Aggregate2025-02-15
Gemma 2 9B

llm

105.0tok/sQ4_K_M8K, Curated Aggregate2025-02-25
Asure 12B

llm

115.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA GeForce RTX 5080 16GB. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 16GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Won't fit
13-14B
Excellent
Tight
Won't fit
32B
Won't fit
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Reported batch-one examples
Embedding throughput7200 emb/s
Whisper transcription95 x RT
SDXL image gen18 img/min

Source

Re-bench: TEI 1.5 + driver R570 tuning.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Aging

Source-linked row with explicit verification status.

Record date: 2026-05-23; 109 days old.

Open primary source

Owner-reported capture

Owner-reported capture via ollama-local-api on 2026-05-28. Hardware identity and outputs are supplied by the owner; this is not independent replication.

Batch=1, 2048 context, 512 generated tokens, consecutive warm-model runs. Generation TPS uses Ollama eval_count/eval_duration.

kumru-2b

2.4B / Q4_K_M

429.62 tok/s

median 443.66 / min 397.42 / max 454.33 / n=5

brooqs-mistral-turkish-v2-latest

7.2B / Q4_0

160.83 tok/s

median 161.11 / min 159.88 / max 161.79 / n=5

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

Search Amazon listings
  • +Use $999 as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.