DeepSeek V3-0324 is the March 24, 2025 refresh — same 685 B / 37 B-active MoE architecture as base V3, retrained with R1-style reasoning distillation and an MIT license (a meaningful upgrade from the original V3's bespoke DeepSeek license). On MATH-500, AIME, GPQA, and LiveCodeBench it lands between base V3 and R1 while still producing direct answers (no <think> trace). VRAM math is the same as base V3: ~1370 GB FP16, ~378 GB Q4_K_M — so 8x H100 or 8x B200 is the floor. For Lagos and Istanbul builders this is a cloud-API model, not a homelab one. See our [MoE architecture](/glossary/moe) glossary entry for why only 37 B of the 685 B is touched per token.
Parameters
685B
37B active (MoE)
Context
128K
tokens
Min VRAM (Q4)
414.4 GB
weights only
Run locally?
NO
needs multi-GPU
//Quick answer
What hardware do I need to run DeepSeek V3-0324?
DeepSeek V3-0324 needs at minimum 414.4 GB of VRAM at Q4_K_M quantization (1507 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the NVIDIA GH200 Grace Hopper (576 GB VRAM, $65,000 MSRP). Community benchmark submissions are open. This model exceeds 48 GB at Q4, so plan for a 2-or-more-GPU split.
TL;DR, what to buy
Recommended GPU
NVIDIA GH200 Grace Hopper
576 GB VRAM · $65,000 MSRP
Min VRAM at Q4_K_M
414.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below
VRAM requirements by quantization
Weights only. Add ~20-30% for KV cache at typical context lengths.