llm workload · 50 runs on record

Gemma 2 9B (Q4)

Google Gemma 2 9B IT, Q4_K_M, batch 1, 8K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.

  • Prompt 2

    Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.

  • Prompt 3

    List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.

  • Prompt 4

    I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?

  • Prompt 5

    Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m gemma-2-9b-it.Q4_K_M.gguf -c 8192 -ngl 999 --seed 42

Single batch. Google Gemma 2 9B IT. 8K context.

Full leaderboard

Every record for Gemma 2 9B.

Top 10 leaderboard

Gemma 2 9B (Q4) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Gemma 2 9B (Q4). 1. 8× NVIDIA H100 SXM5 80GB (DGX H100) at 1,450 tok/s. 2. Cerebras WSE-3 at 1,420 tok/s. 3. Groq LPU Inference Engine at 580 tok/s. 4. NVIDIA B200 192GB at 332 tok/s. 5. NVIDIA GH200 480GB at 248 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 50 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: 8× NVIDIA H100 SXM5 80GB (DGX H100) at $200,000 delivering 1,450 tok/s; Cerebras WSE-3 at $2,000,000 delivering 1,420 tok/s; Groq LPU Inference Engine at $20,000 delivering 580 tok/s; B200 192GB at $39,999 delivering 332 tok/s; GH200 480GB at $45,000 delivering 248 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
8× NVIDIA H100 SXM5 80GB (DGX H100)Datacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
1.4ktok/s
640 GB5600 W$200k0.26 tok/s/W$0.0015/kAmazon
2
Cerebras WSE-3ASIC
Cerebras·FP16·8K ctx
6
1.4ktok/s
44000 GB23000 W$2.00M0.06 tok/s/W$0.0149/kAmazon
3
Groq LPU Inference EngineASIC
Groq·FP8·8K ctx
6
580tok/s
230 GB215 W$20k2.70 tok/s/W$0.36/MAmazon
#4
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
332tok/s
192 GB1000 W$40k0.33 tok/s/W$0.0013/kAmazon
#5
NVIDIA GH200 480GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
248tok/s
144 GB1000 W$45k0.25 tok/s/W$0.0019/kAmazon
#6
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
235tok/s
80 GB700 W$25k0.34 tok/s/W$0.0011/kAmazon
#7
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
220tok/s
141 GB700 W$30k0.31 tok/s/W$0.0014/kAmazon
#8
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
215tok/s
80 GB700 W$25k0.31 tok/s/W$0.0012/kAmazon
#9
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
188tok/s
192 GB750 W$18k0.25 tok/s/W$0.0010/kAmazon
#10
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
188tok/s
80 GB700 W$25k0.27 tok/s/W$0.0014/kAmazon
#11
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
168tok/s
192 GB750 W$15k0.22 tok/s/W$0.94/MAmazon
#12
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
152tok/s
32 GB575 W$2.0k0.26 tok/s/W$0.14/MAmazon
#13
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
152tok/s
32 GB575 W$2.0k0.26 tok/s/W$0.14/MAmazon
#14
NVIDIA A100 SXM4 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
138tok/s
80 GB400 W$15k0.34 tok/s/W$0.0011/kAmazon
#15
2× NVIDIA RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
138tok/s
48 GB900 W$3.2k0.15 tok/s/W$0.24/MAmazon
#16
NVIDIA A100 SXM4 40GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
130tok/s
40 GB400 W$10k0.33 tok/s/W$0.81/MAmazon
#17
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
122tok/s
32 GB575 W$2.0k0.21 tok/s/W$0.17/MAmazon
#18
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
118tok/s
128 GB240 W$3.0k0.49 tok/s/W$0.27/MAmazon
#19
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
118tok/s
48 GB350 W$7.8k0.34 tok/s/W$0.70/MAmazon
#20
AMD Instinct MI250X 128GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
110tok/s
128 GB560 W$12k0.20 tok/s/W$0.0012/kAmazon
#21
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
108tok/s
24 GB450 W$1.6k0.24 tok/s/W$0.16/MAmazon
#22
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
108tok/s
24 GB450 W$1.6k0.24 tok/s/W$0.16/MAmazon
#23
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
105tok/s
16 GB360 W$9990.29 tok/s/W$0.10/MAmazon
#24
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·Q4_K_M·8K ctx
6
104tok/s
48 GB300 W$6.8k0.35 tok/s/W$0.69/MAmazon
#25
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
92tok/s
24 GB450 W$1.6k0.20 tok/s/W$0.18/MAmazon
#26
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
92tok/s
16 GB300 W$7490.31 tok/s/W$0.09/MAmazon
#27
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
86tok/s
16 GB320 W$9990.27 tok/s/W$0.12/MAmazon
#28
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
76tok/s
12 GB250 W$5490.30 tok/s/W$0.08/MAmazon
#29
NVIDIA GeForce RTX 4070 Ti 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
75tok/s
12 GB285 W$7990.26 tok/s/W$0.11/MAmazon
#30
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
72tok/s
12 GB250 W$5490.29 tok/s/W$0.08/MAmazon
#31
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
72tok/s
16 GB320 W$9990.23 tok/s/W$0.15/MAmazon
#32
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·8K ctx
6
72tok/s
24 GB355 W$9990.20 tok/s/W$0.15/MAmazon
#33
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
65tok/s
24 GB350 W$1.5k0.19 tok/s/W$0.24/MAmazon
#34
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
65tok/s
24 GB350 W$1.5k0.19 tok/s/W$0.24/MAmazon
#35
NVIDIA GeForce RTX 4070 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
64tok/s
12 GB200 W$5990.32 tok/s/W$0.10/MAmazon
#36
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
62tok/s
512 GB100 W$10.0k0.62 tok/s/W$0.0017/kAmazon
#37
NVIDIA GeForce RTX 5060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
62tok/s
16 GB180 W$4990.34 tok/s/W$0.09/MAmazon
#38
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
58tok/s
24 GB355 W$9990.16 tok/s/W$0.18/MAmazon
#39
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
48tok/s
128 GB65 W$4.7k0.74 tok/s/W$0.0010/kAmazon
#40
NVIDIA GeForce RTX 4060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
47tok/s
16 GB165 W$4990.28 tok/s/W$0.11/MAmazon
#41
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
44tok/s
128 GB70 W$4.7k0.63 tok/s/W$0.0011/kAmazon
#42
Apple M2 Ultra (76c GPU, 192GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
44tok/s
192 GB80 W$7.0k0.55 tok/s/W$0.0017/kAmazon
#43
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
42tok/s
12 GB190 W$2490.22 tok/s/W$0.06/MAmazon
#44
Apple M3 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
40tok/s
128 GB60 W$4.0k0.67 tok/s/W$0.0011/kAmazon
#45
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·8K ctx
6
35tok/s
12 GB190 W$2490.18 tok/s/W$0.07/MAmazon
#46
NVIDIA GeForce RTX 3060 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
32tok/s
12 GB170 W$3290.19 tok/s/W$0.11/MAmazon
#47
Apple M2 Max (38c GPU, 96GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
32tok/s
96 GB50 W$3.5k0.64 tok/s/W$0.0012/kAmazon
#48
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·8K ctx
6
30tok/s
96 GB120 W$2.2k0.25 tok/s/W$0.78/MAmazon
#49
Apple Mac mini M4 Pro 48GBApple Silicon
Apple·Q4_K_M·8K ctx
6
28tok/s
48 GB35 W$2.0k0.80 tok/s/W$0.76/MAmazon
#50
Apple Mac mini M4 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
14tok/s
24 GB22 W$9990.64 tok/s/W$0.75/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_gemma-2-9b-q4_2026,
  title  = {MyAI Bench: Gemma 2 9B (Q4)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/gemma-2-9b-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Gemma 2 9B (Q4). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/gemma-2-9b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Gemma 2 9B (Q4)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/gemma-2-9b-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, Gemma 2 9B (Q4). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/gemma-2-9b-q4 (accessed 2026-08-27).