llm workload · 22 runs on record

Llama 3 8B (FP16)

Meta Llama 3 8B Instruct, full FP16 weights, batch 1, 2K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.

  • Prompt 2

    Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.

  • Prompt 3

    List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.

  • Prompt 4

    I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?

  • Prompt 5

    Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m llama-3-8b-instruct.fp16.gguf -c 4096 -ngl 999 --seed 42 --prompt-file prompts.txt

Single batch. FP16 weights, ~16GB VRAM. Median of 5 runs, warm-up discarded.

Full leaderboard

Every record for Llama 3 8B FP16.

Top 10 leaderboard

Llama 3 8B (FP16) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Llama 3 8B (FP16). 1. 8× NVIDIA H100 SXM5 80GB at 1,980 tok/s. 2. Cerebras WSE-3 (CS-3) at 1,850 tok/s. 3. Groq LPU (cloud) at 1,850 tok/s. 4. Groq LPU (per chip) at 750 tok/s. 5. NVIDIA B200 192GB at 520 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 22 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: 8× NVIDIA H100 SXM5 80GB at $200,000 delivering 1,980 tok/s; Cerebras WSE-3 (CS-3) at $2,500,000 delivering 1,850 tok/s; Groq LPU (cloud) at $20,000 delivering 1,850 tok/s; Groq LPU (per chip) at $20,000 delivering 750 tok/s; B200 192GB at $39,999 delivering 520 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
8× NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·FP16·8K ctx
6
2.0ktok/s
640 GB5600 W$200k0.35 tok/s/W$0.0011/kAmazon
2
Cerebras WSE-3 (CS-3)ASIC
Cerebras·FP16·8K ctx
6
1.9ktok/s
44 GB23000 W$2.50M0.08 tok/s/W$0.0143/kAmazon
3
Groq LPU (cloud)ASIC
Groq·FP16·8K ctx
6
1.9ktok/s
230 GB350 W$20k5.29 tok/s/W$0.11/MAmazon
#4
Groq LPU (per chip)ASIC
Groq·FP16·8K ctx
6
750tok/s
0.23 GB215 W$20k3.49 tok/s/W$0.28/MAmazon
#5
NVIDIA B200 192GBDatacenter GPU
NVIDIA·FP16·8K ctx
6
520tok/s
192 GB1000 W$40k0.52 tok/s/W$0.81/MAmazon
#6
Google TPU v5pASIC
Google·FP16·8K ctx
6
410tok/s
95 GB700 W$38k0.59 tok/s/W$0.98/MAmazon
#7
NVIDIA GH200 Grace Hopper 480GBDatacenter GPU
NVIDIA·FP16·8K ctx
6
380tok/s
96 GB1000 W$44k0.38 tok/s/W$0.0012/kAmazon
#8
NVIDIA H200 141GBDatacenter GPU
NVIDIA·FP16·4K ctx
6
340tok/s
141 GB700 W$30k0.49 tok/s/W$0.93/MAmazon
#9
AWS Trainium2ASIC
AWS·FP16·8K ctx
6
320tok/s
96 GB500 W$21k0.64 tok/s/W$0.69/MAmazon
#10
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·FP16·4K ctx
5
282tok/s
σ 11.3
80 GB700 W$25k0.40 tok/s/W$0.94/MAmazon
#11
AMD Instinct MI300X 192GBDatacenter GPU
AMD·FP16·8K ctx
6
260tok/s
192 GB750 W$18k0.35 tok/s/W$0.73/MAmazon
#12
Google TPU v5eASIC
Google·FP16·4K ctx
6
240tok/s
16 GB170 W$9.0k1.41 tok/s/W$0.40/MAmazon
#13
NVIDIA A100 80GB SXMDatacenter GPU
NVIDIA·FP16·4K ctx
6
215tok/s
80 GB400 W$15k0.54 tok/s/W$0.74/MAmazon
#14
NVIDIA A100 40GBDatacenter GPU
NVIDIA·FP16·4K ctx
6
195tok/s
40 GB400 W$9.0k0.49 tok/s/W$0.49/MAmazon
#15
AMD Instinct MI250X 128GBDatacenter GPU
AMD·FP16·4K ctx
6
188tok/s
128 GB560 W$14k0.34 tok/s/W$0.79/MAmazon
#16
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·FP16·4K ctx
6
175tok/s
48 GB350 W$7.8k0.50 tok/s/W$0.47/MAmazon
#17
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·FP16·4K ctx
6
96tok/s
32 GB575 W$2.0k0.17 tok/s/W$0.22/MAmazon
#18
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·FP16·8K ctx
6
92tok/s
128 GB240 W$3.0k0.38 tok/s/W$0.34/MAmazon
#19
AMD Radeon Pro W7900 48GBPro GPU
AMD·FP16·4K ctx
6
88tok/s
48 GB295 W$4.0k0.30 tok/s/W$0.48/MAmazon
#20
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·FP16·4K ctx
6
88tok/s
48 GB300 W$6.8k0.29 tok/s/W$0.82/MAmazon
#21
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·FP16·4K ctx
6
88tok/s
16 GB360 W$9990.24 tok/s/W$0.12/MAmazon
#22
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·FP16·4K ctx
6
72tok/s
24 GB450 W$1.6k0.16 tok/s/W$0.23/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_llama3-8b-fp16_2026,
  title  = {MyAI Bench: Llama 3 8B (FP16)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Llama 3 8B (FP16). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16
MLA
MyAIHardware Contributors. "MyAI Bench: Llama 3 8B (FP16)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16. Accessed 2026-08-27.
Plain text
MyAI Bench, Llama 3 8B (FP16). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/llama3-8b-fp16 (accessed 2026-08-27).