NVIDIA GeForce RTX 5090 32GB
8B–32B quantized models with room reserved for runtime and KV cache.
MyAI rating
Not scored
N/A
32 GB
575 W
No benchmark
$2.0k
Real GPUs. Real benchmark coverage. Real next clicks. This page now surfaces your actual AI hardware catalog and sends every serious card into its own benchmark detail page.
GPUs surfaced
43
Consumer, pro, and datacenter cards
Benchmark records
371
Linked back to methodology-aware detail pages
Average rating
Not scored
Composite scoring from throughput, trust, and coverage
Top consumer pick
5090 32GB
Editorial + benchmark-backed
Fresh from your rig
202.10 tok/s on llama3.2:3b / 115.60 tok/s on qwen2.5:7b. Captured on 2026-06-01 with ollama-local-api.
8B–32B quantized models with room reserved for runtime and KV cache.
MyAI rating
Not scored
N/A
32 GB
575 W
No benchmark
$2.0k
Benchmark-backed GPU profile with hardware detail, trust signals, and buying path.
MyAI rating
Not scored
N/A
48 GB
295 W
No benchmark
$4.0k
70B quantized inference. 70B FP16 and 405B Q4 require more memory than one 80GB device.
MyAI rating
Not scored
N/A
80 GB
700 W
No benchmark
$25k
Quick picks
Best local 70B path
Prioritize VRAM headroom and clean 32B-to-70B scaling.
Best value class
The used-market sweet spot still matters for serious builders.
Best scale-up card
When this turns into serving, capacity and interconnect win.
Top LLM throughput
Browse hardwares
Search by name, slice by class, then jump straight into the hardware brief. These cards are built from your benchmark records, not a static mock list.
Showing 43 GPUs / All classes / all brands.
Rating
Not scored
80 GB
700 W
No benchmark
$25k
Top benchmark
Embedding throughput / 12400 emb/s
The default datacenter inference GPU through 2025. Right answer for production multi-tenant serving. Not for homelab — this is enterprise/cloud territory.
Rating
Not scored
141 GB
700 W
No benchmark
$35k
Top benchmark
Embedding throughput / 15200 emb/s
The capacity king. Right answer when 80GB isn't enough and you need 141GB. For everything else, H100 or L40S are more economical.
Rating
Not scored
180 GB
1000 W
No benchmark
$40k
Top benchmark
Embedding throughput / 23500 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
48 GB
350 W
No benchmark
$7.0k
Top benchmark
Embedding throughput / 10500 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
32 GB
575 W
No benchmark
$2.0k
Top benchmark
Embedding throughput / 12200 emb/s
Consider it for a 32GB CUDA workload after checking the artifact and total system cost. Choose more GPU memory for full-residency 70B Q4.
Rating
Not scored
24 GB
450 W
No benchmark
$1.6k
Top benchmark
Embedding throughput / 8500 emb/s
Buy if you want the best single-card experience for 32B-class models and either bought at MSRP or found a clean used unit. Skip if 70B is your daily target or if 16GB cards cover your needs.
Rating
Not scored
16 GB
320 W
No benchmark
$999
Top benchmark
Embedding throughput / 6900 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
12 GB
285 W
No benchmark
$799
Top benchmark
Phi-3 Mini / 165 tok/s
Buy as a gaming-primary, AI-secondary card. For pure AI, the 12GB ceiling is limiting. Consider used 3090 (24GB) or 4060 Ti 16GB instead.
Rating
Not scored
24 GB
350 W
No benchmark
$999
Top benchmark
Embedding throughput / 4900 emb/s
Buy used if you can find a clean unit under $700. The value-for-VRAM proposition is unmatched — nothing else gives you 24GB at this price. Pair with a second used 3090 for 48GB combined VRAM at ~$1,400 total.
Rating
Not scored
12 GB
170 W
No benchmark
$329
Top benchmark
Phi-3 Mini / 92.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
192 GB
750 W
No benchmark
$15k
Top benchmark
Embedding throughput / 9800 emb/s
Best for organizations with existing AMD infrastructure or those specifically chasing $/VRAM ratios at datacenter scale. The software maturity gap vs NVIDIA is real but narrowing.
Rating
Not scored
24 GB
355 W
No benchmark
$899
Top benchmark
Embedding throughput / 5800 emb/s
Buy for Linux AI builds where $/VRAM is the priority and you're comfortable with the ROCm ecosystem. Skip for Windows, for production, or if you value plug-and-play over tinkering.
Your PC result
202.10 tok/s on llama-3.2-3b-instruct
115.60 tok/s on qwen-2.5-7b-instruct / Windows 10 build 26200 + ROCm 7.10 + Ollama 0.24.0 (HIPBLAS). 4 runs per model after a discarded warmup; decode_tps reported as per-run eval_count/eval_duration. Variance <= 0.4% on reproduced models. Captured by site owner on declared rig. Prompt SHA-256: 5e752b688e0761f321a5721d2fa6b57678bf1de1a1180bb4fc9b7e709bafb6a5.
Rating
Not scored
16 GB
263 W
No benchmark
$499
Top benchmark
Phi-3 Mini / 132 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
48 GB
295 W
No benchmark
$4.0k
Top benchmark
Llama 3 8B FP16 / 88.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
12 GB
190 W
No benchmark
$249
Top benchmark
Phi-3 Mini / 86.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
12 GB
200 W
No benchmark
$549
Top benchmark
Phi-3 Mini / 138 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
165 W
No benchmark
$499
Top benchmark
Mistral 7B / 56.0 tok/s
Buy for budget AI builds where 16GB VRAM capacity matters more than token generation speed. Best new-card value for local AI experimentation. Skip if you can stretch to a used 3090 at $600-700.
Rating
Not scored
8 GB
115 W
No benchmark
$299
Top benchmark
Phi-3 Mini / 95.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
48 GB
300 W
No benchmark
$6.8k
Top benchmark
Embedding throughput / 8200 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
48 GB
300 W
No benchmark
$5.5k
Top benchmark
Phi-3 Mini / 162 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
230 W
No benchmark
$1.8k
Top benchmark
Phi-3 Mini / 122 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
140 W
No benchmark
$1.2k
Top benchmark
Phi-3 Mini / 92.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
80 GB
400 W
No benchmark
$15k
Top benchmark
Embedding throughput / 9500 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
40 GB
400 W
No benchmark
$10k
Top benchmark
Llama 3 8B FP16 / 195 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
576 GB
1000 W
No benchmark
$65k
Top benchmark
Llama 3 8B FP16 / 380 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
128 GB
550 W
No benchmark
$8.0k
Top benchmark
Llama 3 8B FP16 / 188 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
64 GB
300 W
No benchmark
$17k
Top benchmark
Phi-3 Mini / 175 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
12 GB
245 W
No benchmark
$399
Top benchmark
Mistral 7B / 58.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
8 GB
165 W
No benchmark
$269
Top benchmark
Phi-3 Mini / 88.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
225 W
No benchmark
$349
Top benchmark
Phi-3 Mini / 72.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
8 GB
225 W
No benchmark
$199
Top benchmark
Phi-3 Mini / 58.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
48 GB
900 W
No benchmark
$3.2k
Top benchmark
Gemma 2 9B / 138 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
96 GB
1400 W
No benchmark
$4.5k
Top benchmark
Qwen 2.5 14B / 78.0 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
640 GB
5600 W
No benchmark
$200k
Top benchmark
Llama 3 8B FP16 / 1980 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
360 W
No benchmark
$999
Top benchmark
Embedding throughput / 8200 emb/s
Buy for mixed gaming+AI or as a secondary GPU in multi-card rigs. Skip if your primary use case is local AI — used 3090 gives you 24GB at lower cost, and 5070 Ti 16GB gives similar VRAM for $250 less.
Your PC result
429.62 tok/s on kumru-2b
160.83 tok/s on brooqs-mistral-turkish-v2-latest / Batch=1, 2048 context, 512 generated tokens, consecutive warm-model runs. Generation TPS uses Ollama eval_count/eval_duration.
Rating
Not scored
16 GB
285 W
No benchmark
$749
Top benchmark
Embedding throughput / 7100 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
12 GB
250 W
No benchmark
$549
Top benchmark
Phi-3 Mini / 165 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
16 GB
180 W
No benchmark
$429
Top benchmark
Phi-3 Mini / 132 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
8 GB
180 W
No benchmark
$379
Top benchmark
Phi-3 Mini / 158 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
8 GB
150 W
No benchmark
$299
Top benchmark
Phi-3 Mini / 145 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
128 GB
240 W
No benchmark
$3.0k
Top benchmark
DeepSeek-R1 7B / 145 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
192 GB
750 W
No benchmark
$15k
Top benchmark
Embedding throughput / 14800 emb/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.
Rating
Not scored
144 GB
1000 W
No benchmark
$45k
Top benchmark
Mistral 7B / 305 tok/s
Open the hardware brief for methodology, confidence ranges, and buyer guidance.