Llama 3 8B (FP16)
Meta Llama 3 8B Instruct, full FP16 weights, batch 1, 2K context.
Primary metric: Tokens / sec (tok/s)
Reference prompts
Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.
- Prompt 1
Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.
- Prompt 2
Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.
- Prompt 3
List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.
- Prompt 4
I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?
- Prompt 5
Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.
Reference runtime command
A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.
llama-server -m llama-3-8b-instruct.fp16.gguf -c 4096 -ngl 999 --seed 42 --prompt-file prompts.txtSingle batch. FP16 weights, ~16GB VRAM. Median of 5 runs, warm-up discarded.
Full leaderboard
Every record for Llama 3 8B FP16.
Top 10 leaderboard
Llama 3 8B (FP16) · sorted by tokens / sec
Value frontier, MSRP vs tokens / sec
Each dot is a device. Top-left is best value (cheap + fast).
Full leaderboard
Click any row for detailed breakdown. Click column headers to sort.
| # | Device | Verif | Buy | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 8× NVIDIA H100 SXM5 80GBDatacenter GPU NVIDIA·FP16·8K ctx | 6 | 2.0ktok/s | 640 GB | 5600 W | $200k | 0.35 tok/s/W | $0.0011/k | Amazon | |
| 2 | Cerebras WSE-3 (CS-3)ASIC Cerebras·FP16·8K ctx | 6 | 1.9ktok/s | 44 GB | 23000 W | $2.50M | 0.08 tok/s/W | $0.0143/k | Amazon | |
| 3 | Groq LPU (cloud)ASIC Groq·FP16·8K ctx | 6 | 1.9ktok/s | 230 GB | 350 W | $20k | 5.29 tok/s/W | $0.11/M | Amazon | |
| #4 | Groq LPU (per chip)ASIC Groq·FP16·8K ctx | 6 | 750tok/s | 0.23 GB | 215 W | $20k | 3.49 tok/s/W | $0.28/M | Amazon | |
| #5 | NVIDIA B200 192GBDatacenter GPU NVIDIA·FP16·8K ctx | 6 | 520tok/s | 192 GB | 1000 W | $40k | 0.52 tok/s/W | $0.81/M | Amazon | |
| #6 | Google TPU v5pASIC Google·FP16·8K ctx | 6 | 410tok/s | 95 GB | 700 W | $38k | 0.59 tok/s/W | $0.98/M | Amazon | |
| #7 | NVIDIA GH200 Grace Hopper 480GBDatacenter GPU NVIDIA·FP16·8K ctx | 6 | 380tok/s | 96 GB | 1000 W | $44k | 0.38 tok/s/W | $0.0012/k | Amazon | |
| #8 | NVIDIA H200 141GBDatacenter GPU NVIDIA·FP16·4K ctx | 6 | 340tok/s | 141 GB | 700 W | $30k | 0.49 tok/s/W | $0.93/M | Amazon | |
| #9 | AWS Trainium2ASIC AWS·FP16·8K ctx | 6 | 320tok/s | 96 GB | 500 W | $21k | 0.64 tok/s/W | $0.69/M | Amazon | |
| #10 | NVIDIA H100 SXM5 80GBDatacenter GPU NVIDIA·FP16·4K ctx | 5 | 282tok/s σ 11.3 | 80 GB | 700 W | $25k | 0.40 tok/s/W | $0.94/M | Amazon | |
| #11 | AMD Instinct MI300X 192GBDatacenter GPU AMD·FP16·8K ctx | 6 | 260tok/s | 192 GB | 750 W | $18k | 0.35 tok/s/W | $0.73/M | Amazon | |
| #12 | Google TPU v5eASIC Google·FP16·4K ctx | 6 | 240tok/s | 16 GB | 170 W | $9.0k | 1.41 tok/s/W | $0.40/M | Amazon | |
| #13 | NVIDIA A100 80GB SXMDatacenter GPU NVIDIA·FP16·4K ctx | 6 | 215tok/s | 80 GB | 400 W | $15k | 0.54 tok/s/W | $0.74/M | Amazon | |
| #14 | NVIDIA A100 40GBDatacenter GPU NVIDIA·FP16·4K ctx | 6 | 195tok/s | 40 GB | 400 W | $9.0k | 0.49 tok/s/W | $0.49/M | Amazon | |
| #15 | AMD Instinct MI250X 128GBDatacenter GPU AMD·FP16·4K ctx | 6 | 188tok/s | 128 GB | 560 W | $14k | 0.34 tok/s/W | $0.79/M | Amazon | |
| #16 | NVIDIA L40S 48GBDatacenter GPU NVIDIA·FP16·4K ctx | 6 | 175tok/s | 48 GB | 350 W | $7.8k | 0.50 tok/s/W | $0.47/M | Amazon | |
| #17 | NVIDIA GeForce RTX 5090 32GBConsumer GPU NVIDIA·FP16·4K ctx | 6 | 96tok/s | 32 GB | 575 W | $2.0k | 0.17 tok/s/W | $0.22/M | Amazon | |
| #18 | NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU NVIDIA·FP16·8K ctx | 6 | 92tok/s | 128 GB | 240 W | $3.0k | 0.38 tok/s/W | $0.34/M | Amazon | |
| #19 | AMD Radeon Pro W7900 48GBPro GPU AMD·FP16·4K ctx | 6 | 88tok/s | 48 GB | 295 W | $4.0k | 0.30 tok/s/W | $0.48/M | Amazon | |
| #20 | NVIDIA RTX 6000 Ada 48GBPro GPU NVIDIA·FP16·4K ctx | 6 | 88tok/s | 48 GB | 300 W | $6.8k | 0.29 tok/s/W | $0.82/M | Amazon | |
| #21 | NVIDIA GeForce RTX 5080 16GBConsumer GPU NVIDIA·FP16·4K ctx | 6 | 88tok/s | 16 GB | 360 W | $999 | 0.24 tok/s/W | $0.12/M | Amazon | |
| #22 | NVIDIA GeForce RTX 4090 24GBConsumer GPU NVIDIA·FP16·4K ctx | 6 | 72tok/s | 24 GB | 450 W | $1.6k | 0.16 tok/s/W | $0.23/M | Amazon |
Cite this benchmark
Use this in your paper, blog post, or comparison table.
@misc{myaihardware_llama3-8b-fp16_2026,
title = {MyAI Bench: Llama 3 8B (FP16)},
author = {{MyAIHardware Contributors}},
year = {2026},
url = {https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16},
note = {Version 1.3, accessed 2026-08-27}
}MyAIHardware Contributors. (2026). MyAI Bench: Llama 3 8B (FP16). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16
MyAIHardware Contributors. "MyAI Bench: Llama 3 8B (FP16)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16. Accessed 2026-08-27.
MyAI Bench, Llama 3 8B (FP16). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16 (accessed 2026-08-27).