Qwen 3Text LLMApache 2.0Apr 2025

Qwen 3 32B

Qwen 3 32B is the largest Qwen 3 that fits in a single 24 GB consumer GPU at Q4_K_M — and on reasoning benchmarks it closes most of the gap to the 72B sibling thanks to the hybrid thinking mode. For a Lagos or Istanbul builder running a single RTX 4090 or 5090, this is the strongest local model you can run without resorting to MoE offload. Apache 2.0, 128K context, full multilingual support. The thinking mode produces noticeably longer outputs than Qwen 2.5 32B, so expect tok/s to feel slower at identical hardware unless you disable it.

Parameters
32B
dense
Context
128K
tokens
Min VRAM (Q4)
19.4 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Qwen 3 32B?

Qwen 3 32B needs at minimum 19.4 GB of VRAM at Q4_K_M quantization (70.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4, 32 GB) (32 GB VRAM, $999 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Qwen 3 32B (Qwen 3, 32B params)As of 2025-04-29

TL;DR, what to buy

Recommended GPU
Apple Mac mini (M4, 32 GB)
32 GB VRAM · $999 MSRP
Min VRAM at Q4_K_M
19.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1670.4 GBReferenceTraining-precision weights
Q8_035.2 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K28.5 GBVery high6-bit, near Q8 quality
Q5_K_M24.3 GBHighStrong middle ground
Q4_K_M19.4 GBBalanced (recommended)Default for local deployments
Q4_019.8 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M15.1 GBLossyWhen VRAM is very tight
Q2_K11.3 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 70.4 GB; Q4_K_M needs 19.4 GB.

At full precision (FP16)

  • Intel Data Center GPU Max 1550128 GB · $2,500
  • Apple Mac Studio (M2 Max, 96 GB)96 GB · $2,999
  • Apple Mac Studio (M4 Max, 128 GB)128 GB · $3,499
  • Apple MacBook Pro 16" (M2 Max, 96 GB)96 GB · $3,899
  • Apple Mac Studio (M3 Ultra, 96 GB)96 GB · $3,999
  • Apple MacBook Pro 16" (M3 Max, 128 GB)128 GB · $4,699
  • Apple MacBook Pro 16" (M4 Max, 128 GB)128 GB · $4,699
  • Apple Mac Studio (M1 Ultra, 128 GB)128 GB · $4,799
  • Apple Mac Studio (M2 Ultra, 128 GB)128 GB · $4,799
  • Apple Mac Studio (M3 Ultra, 256 GB)256 GB · $5,599

At Q4_K_M quantization

  • Intel Arc Pro B6024 GB · $500
  • AMD RX 7900 XT20 GB · $749
  • Apple Mac mini (M4, 24 GB)24 GB · $799
  • AMD RX 7900 XTX24 GB · $899
  • NVIDIA RTX 309024 GB · $999
  • Apple Mac mini (M2, 24 GB)24 GB · $999
  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • NVIDIA RTX 3090 Ti24 GB · $1,099
  • Apple MacBook Air 13" (M4, 24 GB)24 GB · $1,199
  • NVIDIA RTX 4000 Ada Generation20 GB · $1,250
  • AMD Radeon AI PRO R970032 GB · $1,299
  • Apple Mac mini (M4 Pro, 24 GB)24 GB · $1,399

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull qwen3:32b
LM Studio
GGUF format
lms get Qwen/Qwen3-32B-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{qwen-3-32b-2025,
  title={Qwen 3 32B},
  author={Alibaba Cloud Qwen Team},
  year={2025},
  url={https://huggingface.co/Qwen/Qwen3-32B}
}
APA
Alibaba Cloud Qwen Team (2025). Qwen 3 32B [Model card]. Hugging Face. https://huggingface.co/Qwen/Qwen3-32B
Plain text
Qwen 3 32B (Qwen 3, Alibaba Cloud Qwen Team, 2025). Available at https://huggingface.co/Qwen/Qwen3-32B.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models