DeepSeek R1Text LLMMITJan 2025

DeepSeek-R1-Distill-Qwen-32B

This distillation transfers R1's chain-of-thought reasoning into a 32 B dense Qwen-2.5 backbone — and it beats GPT-4o-mini and Claude 3.5 Sonnet on AIME and MATH benchmarks. At Q4_K_M it fits in 24 GB (RTX 4090 / 5090 / 3090) with usable context, making it the most accessible top-tier reasoning model for local users. Drawback: every output is preceded by a long <think> trace, so latency-to-first-useful-token is high.

Parameters
32B
dense
Context
128K
tokens
Min VRAM (Q4)
19.4 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run DeepSeek-R1-Distill-Qwen-32B?

DeepSeek-R1-Distill-Qwen-32B needs at minimum 19.4 GB of VRAM at Q4_K_M quantization (70.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4, 32 GB) (32 GB VRAM, $999 MSRP). Measured throughput hits 620 tok/s on Groq LPU (per chip). This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: DeepSeek-R1-Distill-Qwen-32B (DeepSeek R1, 32B params)As of 2025-01-20

TL;DR, what to buy

Recommended GPU
Apple Mac mini (M4, 32 GB)
32 GB VRAM · $999 MSRP
Min VRAM at Q4_K_M
19.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
620 tok/s
on Groq LPU (per chip)

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1670.4 GBReferenceTraining-precision weights
Q8_035.2 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K28.5 GBVery high6-bit, near Q8 quality
Q5_K_M24.3 GBHighStrong middle ground
Q4_K_M19.4 GBBalanced (recommended)Default for local deployments
Q4_019.8 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M15.1 GBLossyWhen VRAM is very tight
Q2_K11.3 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 70.4 GB; Q4_K_M needs 19.4 GB.

At full precision (FP16)

  • Intel Data Center GPU Max 1550128 GB · $2,500
  • Apple Mac Studio (M2 Max, 96 GB)96 GB · $2,999
  • Apple Mac Studio (M4 Max, 128 GB)128 GB · $3,499
  • Apple MacBook Pro 16" (M2 Max, 96 GB)96 GB · $3,899
  • Apple Mac Studio (M3 Ultra, 96 GB)96 GB · $3,999
  • Apple MacBook Pro 16" (M3 Max, 128 GB)128 GB · $4,699
  • Apple MacBook Pro 16" (M4 Max, 128 GB)128 GB · $4,699
  • Apple Mac Studio (M1 Ultra, 128 GB)128 GB · $4,799
  • Apple Mac Studio (M2 Ultra, 128 GB)128 GB · $4,799
  • Apple Mac Studio (M3 Ultra, 256 GB)256 GB · $5,599

At Q4_K_M quantization

  • Intel Arc Pro B6024 GB · $500
  • AMD RX 7900 XT20 GB · $749
  • Apple Mac mini (M4, 24 GB)24 GB · $799
  • AMD RX 7900 XTX24 GB · $899
  • NVIDIA RTX 309024 GB · $999
  • Apple Mac mini (M2, 24 GB)24 GB · $999
  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • NVIDIA RTX 3090 Ti24 GB · $1,099
  • Apple MacBook Air 13" (M4, 24 GB)24 GB · $1,199
  • NVIDIA RTX 4000 Ada Generation20 GB · $1,250
  • AMD Radeon AI PRO R970032 GB · $1,299
  • Apple Mac mini (M4 Pro, 24 GB)24 GB · $1,399

Community benchmarks

36 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
Groq LPU (per chip)620 tok/sFP168192215W
NVIDIA B200 192GB480 tok/sQ4_K_M81921000W
NVIDIA B200 192GB385 tok/sQ4_K_M81921000W
NVIDIA H200 141GB305 tok/sQ4_K_M8192700W
NVIDIA H100 SXM5 80GB285 tok/sQ4_K_M16384700W
AMD Instinct MI300X 192GB248 tok/sQ4_K_M16384750W
NVIDIA H100 SXM5 80GB245 tok/sQ4_K_M8192700W
NVIDIA H200 141GB232 tok/sQ4_K_M8192700W
AMD Instinct MI300X 192GB215 tok/sQ4_K_M8192750W
NVIDIA H100 SXM5 80GB198 tok/sQ4_K_M8192700W
NVIDIA GeForce RTX 5090 32GB185 tok/sQ4_K_M8192575W
AMD Instinct MI300X 192GB168 tok/sQ4_K_M8192750W
NVIDIA GeForce RTX 5090 32GB165 tok/sQ4_K_M8192575W
NVIDIA DGX Spark (Project DIGITS, 128GB)145 tok/sQ4_K_M8192240W
NVIDIA GeForce RTX 5080 16GB138 tok/sQ4_K_M8192360W
NVIDIA L40S 48GB138 tok/sQ4_K_M8192350W
NVIDIA GeForce RTX 4090 24GB125 tok/sQ4_K_M8192450W
NVIDIA GeForce RTX 5080 16GB122 tok/sQ4_K_M8192360W
NVIDIA GeForce RTX 4090 24GB118 tok/sQ4_K_M8192450W
NVIDIA GeForce RTX 5070 Ti 16GB108 tok/sQ4_K_M8192300W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull deepseek-r1:32b
LM Studio
GGUF format
lms get bartowski/DeepSeek-R1-Distill-Qwen-32B-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{deepseek-r1-distill-qwen-32b-2025,
  title={DeepSeek-R1-Distill-Qwen-32B},
  author={DeepSeek AI},
  year={2025},
  url={https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B}
}
APA
DeepSeek AI (2025). DeepSeek-R1-Distill-Qwen-32B [Model card]. Hugging Face. https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
Plain text
DeepSeek-R1-Distill-Qwen-32B (DeepSeek R1, DeepSeek AI, 2025). Available at https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models