Qwen 3 72B is Alibaba's April 2025 flagship dense release and the first major open model to ship a hybrid thinking mode — one set of weights, two behaviors. Set enable_thinking=true for o1-style chain-of-thought on math and code, or false for fast direct answers. It tops Qwen 2.5 72B on AIME, LiveCodeBench, and BFCL while inheriting Qwen's strong multilingual coverage (119 languages). Apache 2.0 license, identical VRAM footprint to Qwen 2.5 72B. At Q4_K_M it fits across 2x RTX 4090 (48 GB) or 1x H100 80 GB. The reasoning mode roughly doubles output tokens, so plan KV-cache headroom accordingly. See [Q4_K_M quantization](/glossary/q4-k-m) for the math.
Parameters
72B
dense
Context
128K
tokens
Min VRAM (Q4)
43.6 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
//Quick answer
What hardware do I need to run Qwen 3 72B?
Qwen 3 72B needs at minimum 43.6 GB of VRAM at Q4_K_M quantization (158.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4 Pro, 64 GB) (64 GB VRAM, $2,199 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.
TL;DR, what to buy
Recommended GPU
Apple Mac mini (M4 Pro, 64 GB)
64 GB VRAM · $2,199 MSRP
Min VRAM at Q4_K_M
43.6 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below
VRAM requirements by quantization
Weights only. Add ~20-30% for KV cache at typical context lengths.