llm workload · 36 runs on record

DeepSeek-R1 Distill 7B (Q4)

DeepSeek-R1-Distill-Qwen-7B, Q4_K_M, batch 1, 8K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    A train leaves station A at 9:00 going 60 mph. Another leaves station B (180 miles away) at 9:30 going 80 mph toward A. When do they meet? Show your work.

  • Prompt 2

    If a 70B-parameter model uses ~140GB at FP16, what is its likely memory footprint at Q4_K_M? Show the arithmetic.

  • Prompt 3

    Walk through which of these is cheaper for 10M tokens/day at 30 days: H100 cloud at $2/hr vs RTX 4090 owned at $1600 + $0.10/kWh.

  • Prompt 4

    Explain step-by-step why FP8 inference can be 2x faster than FP16 on H100 but not on A100.

  • Prompt 5

    Given a 4-bit quantized Llama 3 70B model at ~40GB, how much VRAM headroom is needed for KV cache at 8K context, batch 1?

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m DeepSeek-R1-Distill-Qwen-7B.Q4_K_M.gguf -c 8192 -ngl 999 --seed 42

Reasoning workload. We use 8K context to give the chain-of-thought tokens room. Single batch.

Full leaderboard

Every record for DeepSeek-R1 7B.

Top 10 leaderboard

DeepSeek-R1 Distill 7B (Q4) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for DeepSeek-R1 Distill 7B (Q4). 1. Groq LPU (per chip) at 620 tok/s. 2. NVIDIA B200 192GB at 480 tok/s. 3. NVIDIA B200 192GB at 385 tok/s. 4. NVIDIA H200 141GB at 305 tok/s. 5. NVIDIA H100 SXM5 80GB at 285 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 36 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: Groq LPU (per chip) at $20,000 delivering 620 tok/s; B200 192GB at $39,999 delivering 480 tok/s; B200 192GB at $39,999 delivering 385 tok/s; H200 141GB at $30,000 delivering 305 tok/s; H100 SXM5 80GB at $25,000 delivering 285 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
Groq LPU (per chip)ASIC
Groq·FP16·8K ctx
6
620tok/s
0.23 GB215 W$20k2.88 tok/s/W$0.34/MAmazon
2
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
480tok/s
192 GB1000 W$40k0.48 tok/s/W$0.88/MAmazon
3
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
385tok/s
192 GB1000 W$40k0.39 tok/s/W$0.0011/kAmazon
#4
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
305tok/s
141 GB700 W$30k0.44 tok/s/W$0.0010/kAmazon
#5
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·16K ctx
6
285tok/s
80 GB700 W$25k0.41 tok/s/W$0.93/MAmazon
#6
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·16K ctx
6
248tok/s
192 GB750 W$15k0.33 tok/s/W$0.64/MAmazon
#7
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
245tok/s
80 GB700 W$25k0.35 tok/s/W$0.0011/kAmazon
#8
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
232tok/s
141 GB700 W$30k0.33 tok/s/W$0.0014/kAmazon
#9
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
215tok/s
192 GB750 W$18k0.29 tok/s/W$0.89/MAmazon
#10
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
198tok/s
80 GB700 W$25k0.28 tok/s/W$0.0013/kAmazon
#11
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
185tok/s
32 GB575 W$2.0k0.32 tok/s/W$0.11/MAmazon
#12
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
168tok/s
192 GB750 W$15k0.22 tok/s/W$0.94/MAmazon
#13
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
165tok/s
32 GB575 W$2.0k0.29 tok/s/W$0.13/MAmazon
#14
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
145tok/s
128 GB240 W$3.0k0.60 tok/s/W$0.22/MAmazon
#15
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
138tok/s
16 GB360 W$9990.38 tok/s/W$0.08/MAmazon
#16
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
138tok/s
48 GB350 W$7.8k0.39 tok/s/W$0.60/MAmazon
#17
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
125tok/s
24 GB450 W$1.6k0.28 tok/s/W$0.14/MAmazon
#18
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
122tok/s
16 GB360 W$9990.34 tok/s/W$0.09/MAmazon
#19
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
118tok/s
24 GB450 W$1.6k0.26 tok/s/W$0.14/MAmazon
#20
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
108tok/s
16 GB300 W$7490.36 tok/s/W$0.07/MAmazon
#21
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
105tok/s
16 GB300 W$7490.35 tok/s/W$0.07/MAmazon
#22
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
98tok/s
16 GB320 W$9990.31 tok/s/W$0.11/MAmazon
#23
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
95tok/s
16 GB320 W$9990.30 tok/s/W$0.11/MAmazon
#24
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
82tok/s
24 GB350 W$1.5k0.23 tok/s/W$0.19/MAmazon
#25
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·8K ctx
6
78tok/s
24 GB355 W$9990.22 tok/s/W$0.14/MAmazon
#26
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
75tok/s
12 GB250 W$5490.30 tok/s/W$0.08/MAmazon
#27
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·32K ctx
6
75tok/s
512 GB80 W$8.5k0.94 tok/s/W$0.0012/kAmazon
#28
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·8K ctx
6
70tok/s
24 GB355 W$9990.20 tok/s/W$0.15/MAmazon
#29
NVIDIA GeForce RTX 5060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
62tok/s
16 GB180 W$4490.34 tok/s/W$0.08/MAmazon
#30
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·16K ctx
6
58tok/s
128 GB65 W$5.0k0.89 tok/s/W$0.91/MAmazon
#31
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·8K ctx
6
52tok/s
12 GB190 W$2490.27 tok/s/W$0.05/MAmazon
#32
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
50tok/s
128 GB65 W$4.7k0.77 tok/s/W$0.99/MAmazon
#33
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
48tok/s
128 GB70 W$4.7k0.69 tok/s/W$0.0010/kAmazon
#34
Apple Mac mini M4 Pro 48GBApple Silicon
Apple·Q4_K_M·16K ctx
6
32tok/s
48 GB35 W$2.0k0.91 tok/s/W$0.66/MAmazon
#35
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·8K ctx
6
24tok/s
96 GB120 W$2.2k0.20 tok/s/W$0.97/MAmazon
#36
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·8K ctx
6
22tok/s
64 GB60 W$2.0k0.37 tok/s/W$0.96/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_deepseek-r1-7b-q4_2026,
  title  = {MyAI Bench: DeepSeek-R1 Distill 7B (Q4)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/deepseek-r1-7b-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: DeepSeek-R1 Distill 7B (Q4). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/deepseek-r1-7b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: DeepSeek-R1 Distill 7B (Q4)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/deepseek-r1-7b-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, DeepSeek-R1 Distill 7B (Q4). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/deepseek-r1-7b-q4 (accessed 2026-08-27).