Qwen 3Text LLMApache 2.0Apr 2025

Qwen 3 72B

Qwen 3 72B is Alibaba's April 2025 flagship dense release and the first major open model to ship a hybrid thinking mode — one set of weights, two behaviors. Set enable_thinking=true for o1-style chain-of-thought on math and code, or false for fast direct answers. It tops Qwen 2.5 72B on AIME, LiveCodeBench, and BFCL while inheriting Qwen's strong multilingual coverage (119 languages). Apache 2.0 license, identical VRAM footprint to Qwen 2.5 72B. At Q4_K_M it fits across 2x RTX 4090 (48 GB) or 1x H100 80 GB. The reasoning mode roughly doubles output tokens, so plan KV-cache headroom accordingly. See [Q4_K_M quantization](/glossary/q4-k-m) for the math.

Parameters
72B
dense
Context
128K
tokens
Min VRAM (Q4)
43.6 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Qwen 3 72B?

Qwen 3 72B needs at minimum 43.6 GB of VRAM at Q4_K_M quantization (158.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4 Pro, 64 GB) (64 GB VRAM, $2,199 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Qwen 3 72B (Qwen 3, 72B params)As of 2025-04-29

TL;DR, what to buy

Recommended GPU
Apple Mac mini (M4 Pro, 64 GB)
64 GB VRAM · $2,199 MSRP
Min VRAM at Q4_K_M
43.6 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP16158.4 GBReferenceTraining-precision weights
Q8_079.2 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K64.2 GBVery high6-bit, near Q8 quality
Q5_K_M54.6 GBHighStrong middle ground
Q4_K_M43.6 GBBalanced (recommended)Default for local deployments
Q4_044.6 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M34.1 GBLossyWhen VRAM is very tight
Q2_K25.3 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 158.4 GB; Q4_K_M needs 43.6 GB.

At full precision (FP16)

  • Apple Mac Studio (M3 Ultra, 256 GB)256 GB · $5,599
  • Apple Mac Studio (M2 Ultra, 192 GB)192 GB · $6,599
  • Apple Mac Studio (M3 Ultra, 512 GB)512 GB · $9,499
  • Apple Mac Pro (M2 Ultra, 192 GB)192 GB · $9,599
  • AMD Instinct MI300X192 GB · $14,999
  • AMD Instinct MI325X256 GB · $18,000
  • AMD Instinct MI355X [VERIFY]288 GB · $25,000
  • NVIDIA B100180 GB · $38,000
  • NVIDIA B200180 GB · $40,000
  • NVIDIA GH200 Grace Hopper576 GB · $65,000

At Q4_K_M quantization

  • Apple Mac mini (M4 Pro, 48 GB)48 GB · $1,799
  • Apple Mac mini (M4 Pro, 64 GB)64 GB · $2,199
  • Apple Mac Studio (M1 Max, 64 GB)64 GB · $2,399
  • Apple Mac Studio (M2 Max, 64 GB)64 GB · $2,399
  • Apple MacBook Pro 14" (M4 Pro, 48 GB)48 GB · $2,399
  • Apple Mac Studio (M4 Max, 64 GB)64 GB · $2,499
  • Apple MacBook Pro 16" (M4 Pro, 48 GB)48 GB · $2,499
  • Intel Data Center GPU Max 1550128 GB · $2,500
  • Apple Mac Studio (M2 Max, 96 GB)96 GB · $2,999
  • Apple Mac Studio (M4 Max, 128 GB)128 GB · $3,499
  • Apple MacBook Pro 16" (M1 Max, 64 GB)64 GB · $3,499
  • Apple MacBook Pro 16" (M3 Max, 64 GB)64 GB · $3,499

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull qwen3:72b
LM Studio
GGUF format
lms get Qwen/Qwen3-72B-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{qwen-3-72b-2025,
  title={Qwen 3 72B},
  author={Alibaba Cloud Qwen Team},
  year={2025},
  url={https://huggingface.co/Qwen/Qwen3-72B}
}
APA
Alibaba Cloud Qwen Team (2025). Qwen 3 72B [Model card]. Hugging Face. https://huggingface.co/Qwen/Qwen3-72B
Plain text
Qwen 3 72B (Qwen 3, Alibaba Cloud Qwen Team, 2025). Available at https://huggingface.co/Qwen/Qwen3-72B.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models