llm workload · 38 runs on record

Llama 3 70B (Q4_K_M)

Meta Llama 3 70B Instruct, Q4_K_M GGUF, batch 1, 4K context, with offloading when VRAM-bound.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    A train leaves station A at 9:00 going 60 mph. Another leaves station B (180 miles away) at 9:30 going 80 mph toward A. When do they meet? Show your work.

  • Prompt 2

    If a 70B-parameter model uses ~140GB at FP16, what is its likely memory footprint at Q4_K_M? Show the arithmetic.

  • Prompt 3

    Walk through which of these is cheaper for 10M tokens/day at 30 days: H100 cloud at $2/hr vs RTX 4090 owned at $1600 + $0.10/kWh.

  • Prompt 4

    Explain step-by-step why FP8 inference can be 2x faster than FP16 on H100 but not on A100.

  • Prompt 5

    Given a 4-bit quantized Llama 3 70B model at ~40GB, how much VRAM headroom is needed for KV cache at 8K context, batch 1?

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m llama-3-70b-instruct.Q4_K_M.gguf -c 4096 -ngl 999 --seed 42

Single batch. Q4_K_M (~40GB). When VRAM-bound we note `ngl` value explicitly. Median of 5 runs.

Full leaderboard

Every record for Llama 3 70B Q4.

Top 10 leaderboard

Llama 3 70B (Q4_K_M) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Llama 3 70B (Q4_K_M). 1. NVIDIA H200 141GB at 540 tok/s. 2. Cerebras WSE-3 (CS-3) at 450 tok/s. 3. 8× NVIDIA H100 SXM5 80GB at 412 tok/s. 4. NVIDIA H100 SXM5 80GB at 410 tok/s. 5. AMD Instinct MI300X 192GB at 380 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 38 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: H200 141GB at $30,000 delivering 540 tok/s; Cerebras WSE-3 (CS-3) at $2,500,000 delivering 450 tok/s; 8× NVIDIA H100 SXM5 80GB at $200,000 delivering 412 tok/s; H100 SXM5 80GB at $25,000 delivering 410 tok/s; Instinct MI300X 192GB at $15,000 delivering 380 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
540tok/s
141 GB700 W$30k0.77 tok/s/W$0.59/MAmazon
2
Cerebras WSE-3 (CS-3)ASIC
Cerebras·FP16·8K ctx
6
450tok/s
44 GB23000 W$2.50M0.02 tok/s/W$0.0587/kAmazon
3
8× NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
412tok/s
640 GB5600 W$200k0.07 tok/s/W$0.0051/kAmazon
#4
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
410tok/s
80 GB700 W$25k0.59 tok/s/W$0.64/MAmazon
#5
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
380tok/s
192 GB750 W$15k0.51 tok/s/W$0.42/MAmazon
#6
Groq LPU (8-chip rack)ASIC
Groq·FP16·8K ctx
6
285tok/s
1.84 GB1720 W$160k0.17 tok/s/W$0.0059/kAmazon
#7
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
138tok/s
σ 12.4
192 GB1000 W$40k0.14 tok/s/W$0.0031/kAmazon
#8
NVIDIA GH200 Grace Hopper 480GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
95tok/s
96 GB1000 W$44k0.10 tok/s/W$0.0049/kAmazon
#9
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
84tok/s
σ 8.9
141 GB700 W$30k0.12 tok/s/W$0.0038/kAmazon
#10
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
78tok/s
32 GB575 W$2.0k0.14 tok/s/W$0.27/MAmazon
#11
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
72tok/s
σ 5.8
192 GB750 W$18k0.10 tok/s/W$0.0026/kAmazon
#12
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
5
66tok/s
σ 7.2
80 GB700 W$25k0.09 tok/s/W$0.0040/kAmazon
#13
AWS Trainium2ASIC
AWS·Q4_K_M·8K ctx
6
58tok/s
96 GB500 W$28k0.12 tok/s/W$0.0051/kAmazon
#14
AMD Instinct MI250X 128GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
52tok/s
128 GB560 W$14k0.09 tok/s/W$0.0028/kAmazon
#15
NVIDIA A100 80GB SXMDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
48tok/s
80 GB400 W$15k0.12 tok/s/W$0.0033/kAmazon
#16
NVIDIA A100 40GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
42tok/s
40 GB400 W$9.0k0.10 tok/s/W$0.0023/kAmazon
#17
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
32tok/s
128 GB240 W$3.0k0.13 tok/s/W$0.99/MAmazon
#18
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·Q4_K_M·8K ctx
6
30tok/s
48 GB300 W$6.8k0.10 tok/s/W$0.0024/kAmazon
#19
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
28tok/s
σ 2.1
32 GB575 W$2.0k0.05 tok/s/W$0.76/MAmazon
#20
AMD Instinct MI210 64GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
28tok/s
64 GB300 W$8.0k0.09 tok/s/W$0.0030/kAmazon
#21
2× NVIDIA RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
28tok/s
48 GB900 W$3.2k0.03 tok/s/W$0.0012/kAmazon
#22
AMD Radeon Pro W7900 48GBPro GPU
AMD·Q4_K_M·4K ctx
6
24tok/s
48 GB295 W$4.0k0.08 tok/s/W$0.0018/kAmazon
#23
4× NVIDIA RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
24tok/s
96 GB1400 W$4.0k0.02 tok/s/W$0.0018/kAmazon
#24
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
22tok/s
48 GB350 W$7.8k0.06 tok/s/W$0.0037/kAmazon
#25
NVIDIA A40 48GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
18tok/s
48 GB300 W$5.5k0.06 tok/s/W$0.0032/kAmazon
#26
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
14tok/s
24 GB450 W$1.6k0.03 tok/s/W$0.0012/kAmazon
#27
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
14tok/s
512 GB270 W$9.5k0.05 tok/s/W$0.0072/kAmazon
#28
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
11tok/s
24 GB355 W$9990.03 tok/s/W$0.96/MAmazon
#29
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
9tok/s
24 GB350 W$1.5k0.03 tok/s/W$0.0018/kAmazon
#30
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
9tok/s
128 GB65 W$5.0k0.13 tok/s/W$0.0062/kAmazon
#31
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
7tok/s
16 GB360 W$9990.02 tok/s/W$0.0015/kAmazon
#32
Apple Mac mini M4 Pro 48GBApple Silicon
Apple·Q4_K_M·4K ctx
6
7tok/s
48 GB35 W$2.0k0.19 tok/s/W$0.0033/kAmazon
#33
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
6tok/s
128 GB70 W$4.7k0.09 tok/s/W$0.0083/kAmazon
#34
Apple M3 Max (40c GPU, 64GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
6tok/s
64 GB65 W$3.5k0.09 tok/s/W$0.0067/kAmazon
#35
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·4K ctx
6
5tok/s
96 GB120 W$2.2k0.04 tok/s/W$0.0046/kAmazon
#36
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·4K ctx
6
4tok/s
64 GB60 W$2.0k0.07 tok/s/W$0.0050/kAmazon
#37
Intel Xeon 6980P (128c Granite Rapids)CPU-only
Intel·Q4_K_M·4K ctx
6
3tok/s
, 500 W$18k0.01 tok/s/W$0.0672/kAmazon
#38
AMD Threadripper PRO 7995WX (96-core)CPU-only
AMD·Q4_K_M·4K ctx
6
2tok/s
, 350 W$10.0k0.01 tok/s/W$0.0440/kAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_llama3-70b-q4_2026,
  title  = {MyAI Bench: Llama 3 70B (Q4_K_M)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Llama 3 70B (Q4_K_M). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Llama 3 70B (Q4_K_M)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, Llama 3 70B (Q4_K_M). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4 (accessed 2026-08-27).