llm workload · 14 runs on record

Llama 3 70B (Q8)

Meta Llama 3 70B Instruct, Q8_0 GGUF, batch 1, 4K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    A train leaves station A at 9:00 going 60 mph. Another leaves station B (180 miles away) at 9:30 going 80 mph toward A. When do they meet? Show your work.

  • Prompt 2

    If a 70B-parameter model uses ~140GB at FP16, what is its likely memory footprint at Q4_K_M? Show the arithmetic.

  • Prompt 3

    Walk through which of these is cheaper for 10M tokens/day at 30 days: H100 cloud at $2/hr vs RTX 4090 owned at $1600 + $0.10/kWh.

  • Prompt 4

    Explain step-by-step why FP8 inference can be 2x faster than FP16 on H100 but not on A100.

  • Prompt 5

    Given a 4-bit quantized Llama 3 70B model at ~40GB, how much VRAM headroom is needed for KV cache at 8K context, batch 1?

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m llama-3-70b-instruct.Q8_0.gguf -c 4096 -ngl 999 --seed 42

Single batch. Q8_0 (~75GB). Often requires multi-GPU or partial CPU offload. Median of 5 runs.

Full leaderboard

Every record for Llama 3 70B Q8.

Top 10 leaderboard

Llama 3 70B (Q8) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Llama 3 70B (Q8). 1. Google TPU v5p at 110 tok/s. 2. NVIDIA B200 192GB at 98 tok/s. 3. AWS Trainium2 at 92 tok/s. 4. Google TPU v5e at 62 tok/s. 5. NVIDIA H200 141GB at 58 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 14 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: Google TPU v5p at $38,000 delivering 110 tok/s; B200 192GB at $39,999 delivering 98 tok/s; AWS Trainium2 at $21,000 delivering 92 tok/s; Google TPU v5e at $9,000 delivering 62 tok/s; H200 141GB at $30,000 delivering 58 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
Google TPU v5pASIC
Google·INT8·8K ctx
6
110tok/s
95 GB700 W$38k0.16 tok/s/W$0.0037/kAmazon
2
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q8_0·8K ctx
6
98tok/s
192 GB1000 W$40k0.10 tok/s/W$0.0043/kAmazon
3
AWS Trainium2ASIC
AWS·INT8·8K ctx
6
92tok/s
96 GB500 W$21k0.18 tok/s/W$0.0024/kAmazon
#4
Google TPU v5eASIC
Google·INT8·4K ctx
6
62tok/s
16 GB170 W$9.0k0.36 tok/s/W$0.0015/kAmazon
#5
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q8_0·8K ctx
6
58tok/s
141 GB700 W$30k0.08 tok/s/W$0.0055/kAmazon
#6
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q8_0·8K ctx
6
48tok/s
192 GB750 W$18k0.06 tok/s/W$0.0040/kAmazon
#7
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q8_0·4K ctx
6
42tok/s
80 GB700 W$25k0.06 tok/s/W$0.0063/kAmazon
#8
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q8_0·8K ctx
6
38tok/s
192 GB750 W$15k0.05 tok/s/W$0.0042/kAmazon
#9
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q8_0·8K ctx
6
22tok/s
128 GB240 W$3.0k0.09 tok/s/W$0.0014/kAmazon
#10
2× NVIDIA RTX 4090 24GBConsumer GPU
NVIDIA·Q8_0·4K ctx
6
18tok/s
48 GB900 W$3.2k0.02 tok/s/W$0.0019/kAmazon
#11
4× NVIDIA RTX 3090 24GBConsumer GPU
NVIDIA·Q8_0·4K ctx
6
14tok/s
96 GB1400 W$4.0k0.01 tok/s/W$0.0030/kAmazon
#12
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q8_0·4K ctx
6
14tok/s
32 GB575 W$2.0k0.02 tok/s/W$0.0015/kAmazon
#13
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q8_0·8K ctx
6
9tok/s
512 GB270 W$9.5k0.03 tok/s/W$0.0112/kAmazon
#14
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q8_0·4K ctx
6
5tok/s
128 GB65 W$5.0k0.08 tok/s/W$0.0102/kAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_llama3-70b-q8_2026,
  title  = {MyAI Bench: Llama 3 70B (Q8)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/llama3-70b-q8},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Llama 3 70B (Q8). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/llama3-70b-q8
MLA
MyAIHardware Contributors. "MyAI Bench: Llama 3 70B (Q8)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/llama3-70b-q8. Accessed 2026-08-27.
Plain text
MyAI Bench, Llama 3 70B (Q8). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/llama3-70b-q8 (accessed 2026-08-27).