llm workload · 98 runs on record

Mistral 7B (Q4)

Mistral 7B Instruct v0.3, Q4_K_M, batch 1, 4K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.

  • Prompt 2

    Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.

  • Prompt 3

    List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.

  • Prompt 4

    I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?

  • Prompt 5

    Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m mistral-7b-instruct-v0.3.Q4_K_M.gguf -c 4096 -ngl 999 --seed 42

Single batch. The Mistral 7B numbers should track Llama 3 8B closely on compute-bound devices.

Full leaderboard

Every record for Mistral 7B.

Top 10 leaderboard

Mistral 7B (Q4) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Mistral 7B (Q4). 1. Cerebras WSE-3 at 1,850 tok/s. 2. Cerebras WSE-3 (CS-3) at 1,620 tok/s. 3. Groq LPU (cloud) at 1,280 tok/s. 4. Groq LPU Inference Engine at 750 tok/s. 5. Groq LPU (per chip) at 540 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 95 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: Cerebras WSE-3 at $2,000,000 delivering 1,850 tok/s; Cerebras WSE-3 (CS-3) at $2,500,000 delivering 1,620 tok/s; Groq LPU (cloud) at $20,000 delivering 1,280 tok/s; Groq LPU Inference Engine at $20,000 delivering 750 tok/s; Groq LPU (per chip) at $20,000 delivering 540 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
Cerebras WSE-3ASIC
Cerebras·FP16·4K ctx
6
1.9ktok/s
44000 GB23000 W$2.00M0.08 tok/s/W$0.0114/kAmazon
2
Cerebras WSE-3 (CS-3)ASIC
Cerebras·FP16·8K ctx
6
1.6ktok/s
44 GB23000 W$2.50M0.07 tok/s/W$0.0163/kAmazon
3
Groq LPU (cloud)ASIC
Groq·FP16·8K ctx
6
1.3ktok/s
230 GB350 W$20k3.66 tok/s/W$0.17/MAmazon
#4
Groq LPU Inference EngineASIC
Groq·FP8·4K ctx
6
750tok/s
230 GB215 W$20k3.49 tok/s/W$0.28/MAmazon
#5
Groq LPU (per chip)ASIC
Groq·FP16·8K ctx
6
540tok/s
0.23 GB215 W$20k2.51 tok/s/W$0.39/MAmazon
#6
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
415tok/s
192 GB1000 W$40k0.41 tok/s/W$0.0010/kAmazon
#7
Google TPU v5p (Trillium)ASIC
Google·INT8·4K ctx
6
410tok/s
95 GB300 W$01.37 tok/s/W$0.00/MAmazon
#8
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
305tok/s
141 GB700 W$30k0.44 tok/s/W$0.0010/kAmazon
#9
NVIDIA GH200 480GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
305tok/s
144 GB1000 W$45k0.30 tok/s/W$0.0016/kAmazon
#10
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·16K ctx
6
295tok/s
80 GB700 W$25k0.42 tok/s/W$0.90/MAmazon
#11
AWS Trainium2 Trn2ASIC
AWS·INT8·4K ctx
6
280tok/s
96 GB350 W$00.80 tok/s/W$0.00/MAmazon
#12
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
268tok/s
192 GB750 W$18k0.36 tok/s/W$0.71/MAmazon
#13
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
268tok/s
141 GB700 W$30k0.38 tok/s/W$0.0012/kAmazon
#14
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
215tok/s
80 GB700 W$25k0.31 tok/s/W$0.0012/kAmazon
#15
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
198tok/s
32 GB575 W$2.0k0.34 tok/s/W$0.11/MAmazon
#16
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
195tok/s
192 GB750 W$15k0.26 tok/s/W$0.81/MAmazon
#17
Google TPU v5eASIC
Google·INT8·4K ctx
6
195tok/s
16 GB170 W$01.15 tok/s/W$0.00/MAmazon
#18
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
182tok/s
32 GB575 W$2.0k0.32 tok/s/W$0.12/MAmazon
#19
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
178tok/s
32 GB575 W$2.0k0.31 tok/s/W$0.12/MAmazon
#20
NVIDIA A100 SXM4 80GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
168tok/s
80 GB400 W$15k0.42 tok/s/W$0.94/MAmazon
#21
NVIDIA A100 SXM4 40GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
162tok/s
40 GB400 W$10k0.41 tok/s/W$0.65/MAmazon
#22
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
145tok/s
24 GB450 W$1.6k0.32 tok/s/W$0.12/MAmazon
#23
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
145tok/s
128 GB240 W$3.0k0.60 tok/s/W$0.22/MAmazon
#24
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
142tok/s
48 GB350 W$7.8k0.41 tok/s/W$0.58/MAmazon
#25
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
138tok/s
24 GB450 W$1.6k0.31 tok/s/W$0.12/MAmazon
#26
AMD Instinct MI250X 128GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
135tok/s
128 GB560 W$12k0.24 tok/s/W$0.94/MAmazon
#27
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
128tok/s
24 GB450 W$1.6k0.28 tok/s/W$0.13/MAmazon
#28
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
125tok/s
48 GB300 W$6.8k0.42 tok/s/W$0.57/MAmazon
#29
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
122tok/s
16 GB360 W$9990.34 tok/s/W$0.09/MAmazon
#30
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
118tok/s
16 GB320 W$9990.37 tok/s/W$0.09/MAmazon
#31
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
108tok/s
16 GB320 W$9990.34 tok/s/W$0.10/MAmazon
#32
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
108tok/s
16 GB300 W$7490.36 tok/s/W$0.07/MAmazon
#33
AMD Instinct MI210 64GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
105tok/s
64 GB300 W$7.5k0.35 tok/s/W$0.76/MAmazon
#34
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
102tok/s
16 GB320 W$9990.32 tok/s/W$0.10/MAmazon
#35
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
95tok/s
24 GB355 W$9990.27 tok/s/W$0.11/MAmazon
#36
NVIDIA GeForce RTX 4070 Ti 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
89tok/s
12 GB285 W$7990.31 tok/s/W$0.10/MAmazon
#37
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
88tok/s
24 GB355 W$9990.25 tok/s/W$0.12/MAmazon
#38
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
88tok/s
12 GB250 W$5490.35 tok/s/W$0.07/MAmazon
#39
NVIDIA A40 48GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
88tok/s
48 GB300 W$5.5k0.29 tok/s/W$0.66/MAmazon
#40
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
85tok/s
12 GB250 W$5490.34 tok/s/W$0.07/MAmazon
#41
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
82tok/s
512 GB100 W$10.0k0.82 tok/s/W$0.0013/kAmazon
#42
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
78tok/s
24 GB350 W$1.5k0.22 tok/s/W$0.20/MAmazon
#43
NVIDIA GeForce RTX 4070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
76tok/s
12 GB200 W$5990.38 tok/s/W$0.08/MAmazon
#44
NVIDIA GeForce RTX 5060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
74tok/s
16 GB180 W$4990.41 tok/s/W$0.07/MAmazon
#45
AMD Radeon RX 7800 XT 16GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
72tok/s
16 GB263 W$4990.27 tok/s/W$0.07/MAmazon
#46
NVIDIA GeForce RTX 5060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
68tok/s
16 GB180 W$4490.38 tok/s/W$0.07/MAmazon
#47
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·16K ctx
6
68tok/s
512 GB80 W$8.5k0.85 tok/s/W$0.0013/kAmazon
#48
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
64tok/s
128 GB65 W$4.7k0.98 tok/s/W$0.78/MAmazon
#49
NVIDIA GeForce RTX 5060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
64tok/s
8 GB145 W$2990.44 tok/s/W$0.05/MAmazon
#50
NVIDIA RTX A5000 24GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
64tok/s
24 GB230 W$2.5k0.28 tok/s/W$0.41/MAmazon
#51
Apple M2 Ultra (76c GPU, 192GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
60tok/s
192 GB80 W$7.0k0.75 tok/s/W$0.0012/kAmazon
#52
AMD Radeon RX 7700 XT 12GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
58tok/s
12 GB245 W$4490.24 tok/s/W$0.08/MAmazon
#53
AMD Radeon RX 7800 XT 16GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
58tok/s
16 GB263 W$4990.22 tok/s/W$0.09/MAmazon
#54
NVIDIA GeForce RTX 4060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
56tok/s
16 GB165 W$4990.34 tok/s/W$0.09/MAmazon
#55
NVIDIA GeForce RTX 4060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
56tok/s
16 GB165 W$4990.34 tok/s/W$0.09/MAmazon
#56
Apple M3 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
52tok/s
128 GB60 W$4.0k0.87 tok/s/W$0.81/MAmazon
#57
NVIDIA GeForce RTX 3060 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
48tok/s
12 GB170 W$3290.28 tok/s/W$0.07/MAmazon
#58
NVIDIA GeForce RTX 4060 8GBConsumer GPU
NVIDIA·Q4_K_M·2K ctx
6
48tok/s
8 GB115 W$2990.42 tok/s/W$0.07/MAmazon
#59
AMD Radeon RX 7700 XT 12GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
48tok/s
12 GB245 W$4490.20 tok/s/W$0.10/MAmazon
#60
NVIDIA RTX A4000 16GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
48tok/s
16 GB140 W$1.2k0.34 tok/s/W$0.26/MAmazon
#61
Apple M2 Max (38c GPU, 96GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
42tok/s
96 GB60 W$3.3k0.70 tok/s/W$0.83/MAmazon
#62
Apple M4 Pro (20c GPU, 48GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
42tok/s
48 GB50 W$2.5k0.84 tok/s/W$0.63/MAmazon
#63
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
42tok/s
12 GB190 W$2490.22 tok/s/W$0.06/MAmazon
#64
NVIDIA GeForce RTX 4060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
42tok/s
8 GB115 W$2990.36 tok/s/W$0.07/MAmazon
#65
Apple M2 Max (38c GPU, 96GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
42tok/s
96 GB50 W$3.5k0.84 tok/s/W$0.88/MAmazon
#66
Intel Arc A770 16GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
40tok/s
16 GB225 W$3290.18 tok/s/W$0.09/MAmazon
#67
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·4K ctx
6
38tok/s
96 GB120 W$2.2k0.32 tok/s/W$0.61/MAmazon
#68
NVIDIA GeForce RTX 3060 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
38tok/s
12 GB170 W$3290.22 tok/s/W$0.09/MAmazon
#69
Apple Mac mini M4 Pro 48GBApple Silicon
Apple·Q4_K_M·4K ctx
6
38tok/s
48 GB35 W$2.0k1.09 tok/s/W$0.56/MAmazon
#70
Apple Mac mini M4 Pro 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
36tok/s
24 GB35 W$1.4k1.03 tok/s/W$0.41/MAmazon
#71
Apple M4 Pro (20c GPU, 48GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
36tok/s
48 GB35 W$2.0k1.03 tok/s/W$0.59/MAmazon
#72
Apple Mac mini M4 Pro 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
35tok/s
24 GB35 W$1.4k1.00 tok/s/W$0.42/MAmazon
#73
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·8K ctx
6
32tok/s
96 GB120 W$2.2k0.27 tok/s/W$0.73/MAmazon
#74
Intel Arc A770 16GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
32tok/s
16 GB225 W$3490.14 tok/s/W$0.12/MAmazon
#75
Apple M3 Pro (18c GPU, 36GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
28tok/s
36 GB30 W$2.5k0.93 tok/s/W$0.94/MAmazon
#76
Apple M3 Pro (18c GPU, 36GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
25tok/s
36 GB50 W$2.5k0.50 tok/s/W$0.0011/kAmazon
#77
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·4K ctx
6
25tok/s
64 GB60 W$2.0k0.42 tok/s/W$0.85/MAmazon
#78
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·4K ctx
6
24tok/s
64 GB60 W$2.0k0.40 tok/s/W$0.88/MAmazon
#79
Apple M2 Pro (19c GPU, 32GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
24tok/s
32 GB28 W$2.0k0.86 tok/s/W$0.88/MAmazon
#80
Apple Mac mini M4 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
23tok/s
24 GB22 W$7991.04 tok/s/W$0.37/MAmazon
#81
Qualcomm Snapdragon X Elite X1E-84-100NPU
Qualcomm·Q4_K_M·4K ctx
6
22tok/s
, 23 W$1.2k0.96 tok/s/W$0.58/MAmazon
#82
Apple M4 (10c GPU, 24GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
22tok/s
24 GB22 W$9991.00 tok/s/W$0.48/MAmazon
#83
Apple Mac mini M4 16GBApple Silicon
Apple·Q4_K_M·4K ctx
6
20tok/s
16 GB22 W$5990.91 tok/s/W$0.32/MAmazon
#84
Apple Mac mini M4 16GBApple Silicon
Apple·Q4_K_M·4K ctx
6
18tok/s
16 GB18 W$5991.00 tok/s/W$0.35/MAmazon
#85
Intel Xeon 6980P (128 cores)CPU-only
Intel·Q4_K_M·4K ctx
6
18tok/s
, 500 W$18k0.04 tok/s/W$0.0105/kAmazon
#86
AMD Ryzen AI 9 HX 370 (Strix Point NPU 50 TOPS)NPU
AMD·Q4_K_M·4K ctx
6
16tok/s
, 28 W$1.5k0.57 tok/s/W$0.99/MAmazon
#87
Intel Xeon 6980P (128c Granite Rapids)CPU-only
Intel·Q4_K_M·4K ctx
6
15tok/s
, 500 W$18k0.03 tok/s/W$0.0125/kAmazon
#88
AMD Threadripper PRO 7995WX (96-core)CPU-only
AMD·Q4_K_M·4K ctx
6
14tok/s
, 350 W$10.0k0.04 tok/s/W$0.0075/kAmazon
#89
AMD Threadripper Pro 7995WX (96 cores)CPU-only
AMD·Q4_K_M·4K ctx
6
14tok/s
, 350 W$10.0k0.04 tok/s/W$0.0075/kAmazon
#90
Intel Core Ultra 258V (Lunar Lake NPU 48 TOPS)NPU
Intel·Q4_K_M·4K ctx
6
13tok/s
, 17 W$1.3k0.77 tok/s/W$0.0011/kAmazon
#91
Qualcomm Snapdragon X Elite (X1E-84-100)NPU
Qualcomm·INT4·2K ctx
6
12tok/s
, 23 W$1.2k0.52 tok/s/W$0.0011/kAmazon
#92
NVIDIA Jetson Orin NX 16GBEdge
NVIDIA·Q4_K_M·4K ctx
6
10tok/s
16 GB25 W$6990.40 tok/s/W$0.74/MAmazon
#93
Intel Core Ultra 9 285KCPU-only
Intel·Q4_K_M·4K ctx
6
9tok/s
, 250 W$5890.04 tok/s/W$0.69/MAmazon
#94
AMD Ryzen AI 9 HX 370 NPUNPU
AMD·INT4·2K ctx
6
8tok/s
, 28 W$1.5k0.29 tok/s/W$0.0020/kAmazon
#95
NVIDIA Jetson Orin Nano Super 8GBEdge
NVIDIA·Q4_K_M·4K ctx
6
8tok/s
8 GB25 W$2490.32 tok/s/W$0.33/MAmazon
#96
Intel Core Ultra 9 285K NPUNPU
Intel·INT4·2K ctx
6
7tok/s
, 15 W$5890.47 tok/s/W$0.89/MAmazon
#97
NVIDIA Jetson Orin Nano 8GBEdge
NVIDIA·Q4_K_M·4K ctx
6
6tok/s
8 GB15 W$4990.37 tok/s/W$0.96/MAmazon
#98
Raspberry Pi 5 8GBEdge
Raspberry Pi·Q4_K_M·2K ctx
6
2tok/s
, 12 W$800.15 tok/s/W$0.47/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_mistral-7b-q4_2026,
  title  = {MyAI Bench: Mistral 7B (Q4)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/mistral-7b-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Mistral 7B (Q4). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/mistral-7b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Mistral 7B (Q4)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/mistral-7b-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, Mistral 7B (Q4). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/mistral-7b-q4 (accessed 2026-08-27).