Qwen 3 32B is the largest Qwen 3 that fits in a single 24 GB consumer GPU at Q4_K_M — and on reasoning benchmarks it closes most of the gap to the 72B sibling thanks to the hybrid thinking mode. For a Lagos or Istanbul builder running a single RTX 4090 or 5090, this is the strongest local model you can run without resorting to MoE offload. Apache 2.0, 128K context, full multilingual support. The thinking mode produces noticeably longer outputs than Qwen 2.5 32B, so expect tok/s to feel slower at identical hardware unless you disable it.
Parameters
32B
dense
Context
128K
tokens
Min VRAM (Q4)
19.4 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
//Quick answer
What hardware do I need to run Qwen 3 32B?
Qwen 3 32B needs at minimum 19.4 GB of VRAM at Q4_K_M quantization (70.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4, 32 GB) (32 GB VRAM, $999 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.
TL;DR, what to buy
Recommended GPU
Apple Mac mini (M4, 32 GB)
32 GB VRAM · $999 MSRP
Min VRAM at Q4_K_M
19.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below
VRAM requirements by quantization
Weights only. Add ~20-30% for KV cache at typical context lengths.