llm workload · 68 runs on record

Phi-3 Mini 3.8B (Q4)

Microsoft Phi-3 Mini 3.8B Instruct, Q4_K_M, batch 1, 4K context.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.

  • Prompt 2

    Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.

  • Prompt 3

    List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.

  • Prompt 4

    I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?

  • Prompt 5

    Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-server -m phi-3-mini-instruct.Q4_K_M.gguf -c 4096 -ngl 999 --seed 42

Single batch. Tiny 3.8B model — primarily a memory-bandwidth probe.

Full leaderboard

Every record for Phi-3 Mini.

Top 10 leaderboard

Phi-3 Mini 3.8B (Q4) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Phi-3 Mini 3.8B (Q4). 1. NVIDIA H100 SXM5 80GB at 612 tok/s. 2. NVIDIA GeForce RTX 5090 32GB at 348 tok/s. 3. NVIDIA GeForce RTX 5090 32GB at 320 tok/s. 4. Google TPU v5e at 320 tok/s. 5. NVIDIA GeForce RTX 4090 24GB at 285 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 67 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: H100 SXM5 80GB at $25,000 delivering 612 tok/s; GeForce RTX 5090 32GB at $1,999 delivering 348 tok/s; GeForce RTX 5090 32GB at $1,999 delivering 320 tok/s; GeForce RTX 4090 24GB at $1,599 delivering 285 tok/s; GeForce RTX 4090 24GB at $1,599 delivering 245 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·4K ctx
6
612tok/s
80 GB700 W$25k0.87 tok/s/W$0.43/MAmazon
2
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
348tok/s
32 GB575 W$2.0k0.60 tok/s/W$0.06/MAmazon
3
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
320tok/s
32 GB575 W$2.0k0.56 tok/s/W$0.07/MAmazon
#4
Google TPU v5eASIC
Google·INT8·4K ctx
6
320tok/s
16 GB170 W$01.88 tok/s/W$0.00/MAmazon
#5
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
285tok/s
24 GB450 W$1.6k0.63 tok/s/W$0.06/MAmazon
#6
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
245tok/s
24 GB450 W$1.6k0.54 tok/s/W$0.07/MAmazon
#7
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
215tok/s
16 GB320 W$9990.67 tok/s/W$0.05/MAmazon
#8
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
195tok/s
24 GB355 W$9990.55 tok/s/W$0.05/MAmazon
#9
AMD Instinct MI210 64GBDatacenter GPU
AMD·Q4_K_M·4K ctx
6
175tok/s
64 GB300 W$7.5k0.58 tok/s/W$0.45/MAmazon
#10
NVIDIA GeForce RTX 4070 Ti 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
165tok/s
12 GB285 W$7990.58 tok/s/W$0.05/MAmazon
#11
NVIDIA GeForce RTX 5070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
165tok/s
12 GB250 W$5490.66 tok/s/W$0.03/MAmazon
#12
NVIDIA A40 48GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
162tok/s
48 GB300 W$5.5k0.54 tok/s/W$0.36/MAmazon
#13
NVIDIA GeForce RTX 5060 Ti 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
158tok/s
8 GB180 W$3790.88 tok/s/W$0.03/MAmazon
#14
NVIDIA GeForce RTX 5060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
145tok/s
8 GB145 W$2991.00 tok/s/W$0.02/MAmazon
#15
NVIDIA GeForce RTX 4070 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
138tok/s
12 GB200 W$5990.69 tok/s/W$0.05/MAmazon
#16
NVIDIA GeForce RTX 5060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
132tok/s
16 GB180 W$4490.73 tok/s/W$0.04/MAmazon
#17
AMD Radeon RX 7800 XT 16GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
132tok/s
16 GB263 W$4990.50 tok/s/W$0.04/MAmazon
#18
NVIDIA RTX A5000 24GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
122tok/s
24 GB230 W$2.5k0.53 tok/s/W$0.22/MAmazon
#19
NVIDIA GeForce RTX 5060 Ti 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
118tok/s
8 GB165 W$3790.71 tok/s/W$0.03/MAmazon
#20
NVIDIA GeForce RTX 5060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
98tok/s
8 GB145 W$2990.68 tok/s/W$0.03/MAmazon
#21
NVIDIA GeForce RTX 4060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
95tok/s
8 GB115 W$2990.83 tok/s/W$0.03/MAmazon
#22
NVIDIA GeForce RTX 3060 12GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
92tok/s
12 GB170 W$3290.54 tok/s/W$0.04/MAmazon
#23
NVIDIA RTX A4000 16GBPro GPU
NVIDIA·Q4_K_M·4K ctx
6
92tok/s
16 GB140 W$1.2k0.66 tok/s/W$0.14/MAmazon
#24
NVIDIA GeForce RTX 4060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
88tok/s
8 GB115 W$2990.77 tok/s/W$0.04/MAmazon
#25
AMD Radeon RX 7600 8GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
88tok/s
8 GB165 W$2690.53 tok/s/W$0.03/MAmazon
#26
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
86tok/s
12 GB190 W$2490.45 tok/s/W$0.03/MAmazon
#27
NVIDIA GeForce RTX 4060 8GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
6
84tok/s
8 GB115 W$2990.73 tok/s/W$0.04/MAmazon
#28
AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)NPU
AMD·Q4_K_M·4K ctx
6
78tok/s
96 GB120 W$2.2k0.65 tok/s/W$0.30/MAmazon
#29
Apple M4 Pro (20c GPU, 48GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
78tok/s
48 GB35 W$2.0k2.23 tok/s/W$0.27/MAmazon
#30
Intel Arc B580 12GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
76tok/s
12 GB190 W$2490.40 tok/s/W$0.03/MAmazon
#31
Apple Mac mini M4 Pro 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
72tok/s
24 GB35 W$1.4k2.06 tok/s/W$0.20/MAmazon
#32
Intel Arc A770 16GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
72tok/s
16 GB225 W$3490.32 tok/s/W$0.05/MAmazon
#33
Apple M3 Pro (18c GPU, 36GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
62tok/s
36 GB30 W$2.5k2.07 tok/s/W$0.43/MAmazon
#34
AMD Radeon RX 7600 8GBConsumer GPU
AMD·Q4_K_M·4K ctx
6
58tok/s
8 GB165 W$2690.35 tok/s/W$0.05/MAmazon
#35
Intel Arc A750 8GBConsumer GPU
Intel·Q4_K_M·4K ctx
6
58tok/s
8 GB225 W$2490.26 tok/s/W$0.05/MAmazon
#36
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·4K ctx
6
55tok/s
64 GB60 W$2.0k0.92 tok/s/W$0.38/MAmazon
#37
Apple M2 Pro (19c GPU, 32GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
54tok/s
32 GB28 W$2.0k1.93 tok/s/W$0.39/MAmazon
#38
Apple Mac mini M4 24GBApple Silicon
Apple·Q4_K_M·4K ctx
6
52tok/s
24 GB22 W$7992.36 tok/s/W$0.16/MAmazon
#39
NVIDIA Jetson AGX Orin 64GBEdge
NVIDIA·Q4_K_M·4K ctx
6
48tok/s
64 GB60 W$2.0k0.80 tok/s/W$0.44/MAmazon
#40
Apple M4 (10c GPU, 24GB)Apple Silicon
Apple·Q4_K_M·4K ctx
6
48tok/s
24 GB22 W$9992.18 tok/s/W$0.22/MAmazon
#41
Apple Mac mini M4 16GBApple Silicon
Apple·Q4_K_M·4K ctx
6
45tok/s
16 GB22 W$5992.04 tok/s/W$0.14/MAmazon
#42
Apple Mac mini M4 16GBApple Silicon
Apple·Q4_K_M·4K ctx
6
42tok/s
16 GB18 W$5992.33 tok/s/W$0.15/MAmazon
#43
Intel Xeon 6980P (128 cores)CPU-only
Intel·Q4_K_M·4K ctx
6
42tok/s
, 500 W$18k0.08 tok/s/W$0.0045/kAmazon
#44
NVIDIA Jetson Orin NX 16GBEdge
NVIDIA·Q4_K_M·4K ctx
6
38tok/s
16 GB25 W$5991.52 tok/s/W$0.17/MAmazon
#45
Qualcomm Snapdragon X Elite X1E-84-100NPU
Qualcomm·Q4_K_M·4K ctx
6
38tok/s
, 23 W$1.2k1.65 tok/s/W$0.33/MAmazon
#46
AMD Threadripper Pro 7995WX (96 cores)CPU-only
AMD·Q4_K_M·4K ctx
6
36tok/s
, 350 W$10.0k0.10 tok/s/W$0.0029/kAmazon
#47
NVIDIA Jetson Orin Nano Super 8GBEdge
NVIDIA·Q4_K_M·4K ctx
6
32tok/s
8 GB25 W$2491.28 tok/s/W$0.08/MAmazon
#48
Intel Xeon 6980P (128c Granite Rapids)CPU-only
Intel·Q4_K_M·4K ctx
6
32tok/s
, 500 W$18k0.06 tok/s/W$0.0059/kAmazon
#49
AMD Ryzen AI 9 HX 370 (Strix Point NPU 50 TOPS)NPU
AMD·Q4_K_M·4K ctx
6
32tok/s
, 28 W$1.5k1.14 tok/s/W$0.49/MAmazon
#50
Intel Core Ultra 9 288V (Lunar Lake)NPU
Intel·Q4_K_M·4K ctx
6
30tok/s
, 17 W$1.4k1.76 tok/s/W$0.49/MAmazon
#51
AMD Threadripper PRO 7995WX (96-core)CPU-only
AMD·Q4_K_M·4K ctx
6
28tok/s
, 350 W$10.0k0.08 tok/s/W$0.0038/kAmazon
#52
Qualcomm Snapdragon X Elite (X1E-84-100)NPU
Qualcomm·INT4·4K ctx
6
28tok/s
, 23 W$1.2k1.22 tok/s/W$0.45/MAmazon
#53
Intel Core Ultra 258V (Lunar Lake NPU 48 TOPS)NPU
Intel·Q4_K_M·4K ctx
6
28tok/s
, 17 W$1.3k1.65 tok/s/W$0.49/MAmazon
#54
AMD Ryzen AI 9 365 (Strix Point)NPU
AMD·Q4_K_M·4K ctx
6
28tok/s
, 28 W$1.3k1.00 tok/s/W$0.49/MAmazon
#55
Qualcomm Snapdragon X Elite (X1E-84-100)NPU
Qualcomm·INT4·2K ctx
6
25tok/s
, 23 W$1.2k1.09 tok/s/W$0.51/MAmazon
#56
Intel Core Ultra 9 285KCPU-only
Intel·Q4_K_M·4K ctx
6
24tok/s
, 250 W$5890.10 tok/s/W$0.26/MAmazon
#57
Intel Core Ultra 9 288V (Lunar Lake) NPUNPU
Intel·INT4·2K ctx
6
22tok/s
, 17 W$5491.29 tok/s/W$0.26/MAmazon
#58
Intel Core Ultra 7 258V (Lunar Lake)NPU
Intel·INT4·2K ctx
6
22tok/s
, 17 W$1.3k1.29 tok/s/W$0.62/MAmazon
#59
NVIDIA Jetson Orin NX 16GBEdge
NVIDIA·Q4_K_M·4K ctx
6
22tok/s
16 GB25 W$6990.88 tok/s/W$0.34/MAmazon
#60
NVIDIA Jetson Orin Nano Super 8GBEdge
NVIDIA·Q4_K_M·4K ctx
6
21tok/s
8 GB25 W$2490.84 tok/s/W$0.13/MAmazon
#61
AMD Ryzen AI 9 HX 370 NPUNPU
AMD·INT4·2K ctx
6
19tok/s
, 28 W$1.5k0.68 tok/s/W$0.83/MAmazon
#62
NVIDIA Jetson Orin Nano 8GBEdge
NVIDIA·Q4_K_M·4K ctx
6
18tok/s
8 GB15 W$4991.20 tok/s/W$0.29/MAmazon
#63
AMD Ryzen AI 9 365 NPUNPU
AMD·INT4·2K ctx
6
16tok/s
, 28 W$1.2k0.57 tok/s/W$0.79/MAmazon
#64
Intel Core Ultra 9 285K NPUNPU
Intel·INT4·2K ctx
6
14tok/s
, 15 W$5890.93 tok/s/W$0.44/MAmazon
#65
Raspberry Pi AI Hat+ (Hailo-8L 13 TOPS)Edge
Raspberry Pi·INT8·4K ctx
6
6tok/s
, 8 W$700.75 tok/s/W$0.12/MAmazon
#66
Raspberry Pi 5 + AI HAT+ (Hailo-8)Edge
Raspberry Pi·INT8·2K ctx
6
4tok/s
8 GB12 W$1500.35 tok/s/W$0.38/MAmazon
#67
Raspberry Pi 5 8GBEdge
Raspberry Pi·Q4_K_M·2K ctx
6
4tok/s
, 12 W$800.35 tok/s/W$0.20/MAmazon
#68
Raspberry Pi 5 (8GB)Edge
Raspberry Pi·Q4_K_M·4K ctx
6
4tok/s
, 12 W$800.29 tok/s/W$0.24/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_phi3-mini-q4_2026,
  title  = {MyAI Bench: Phi-3 Mini 3.8B (Q4)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/phi3-mini-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Phi-3 Mini 3.8B (Q4). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/phi3-mini-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Phi-3 Mini 3.8B (Q4)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/phi3-mini-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, Phi-3 Mini 3.8B (Q4). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/phi3-mini-q4 (accessed 2026-08-27).