DeepSeek V3Text LLMMITMar 2025

DeepSeek V3-0324

DeepSeek V3-0324 is the March 24, 2025 refresh — same 685 B / 37 B-active MoE architecture as base V3, retrained with R1-style reasoning distillation and an MIT license (a meaningful upgrade from the original V3's bespoke DeepSeek license). On MATH-500, AIME, GPQA, and LiveCodeBench it lands between base V3 and R1 while still producing direct answers (no <think> trace). VRAM math is the same as base V3: ~1370 GB FP16, ~378 GB Q4_K_M — so 8x H100 or 8x B200 is the floor. For Lagos and Istanbul builders this is a cloud-API model, not a homelab one. See our [MoE architecture](/glossary/moe) glossary entry for why only 37 B of the 685 B is touched per token.

Parameters
685B
37B active (MoE)
Context
128K
tokens
Min VRAM (Q4)
414.4 GB
weights only
Run locally?
NO
needs multi-GPU
Quick answer

What hardware do I need to run DeepSeek V3-0324?

DeepSeek V3-0324 needs at minimum 414.4 GB of VRAM at Q4_K_M quantization (1507 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the NVIDIA GH200 Grace Hopper (576 GB VRAM, $65,000 MSRP). Community benchmark submissions are open. This model exceeds 48 GB at Q4, so plan for a 2-or-more-GPU split.

Source: MyAIHardware model card: DeepSeek V3-0324 (DeepSeek V3, 685B params)As of 2025-03-24

TL;DR, what to buy

Recommended GPU
NVIDIA GH200 Grace Hopper
576 GB VRAM · $65,000 MSRP
Min VRAM at Q4_K_M
414.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP161507.0 GBReferenceTraining-precision weights
Q8_0753.5 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K610.3 GBVery high6-bit, near Q8 quality
Q5_K_M519.9 GBHighStrong middle ground
Q4_K_M414.4 GBBalanced (recommended)Default for local deployments
Q4_0423.8 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M324.0 GBLossyWhen VRAM is very tight
Q2_K241.1 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 1507.0 GB; Q4_K_M needs 414.4 GB.

At full precision (FP16)

No GPU in our database has enough VRAM for FP16. Use Q4_K_M or multi-GPU.

At Q4_K_M quantization

  • Apple Mac Studio (M3 Ultra, 512 GB)512 GB · $9,499
  • NVIDIA GH200 Grace Hopper576 GB · $65,000

Multi-GPU splits (Q4_K_M target: 414 GB)

GPUPer-card VRAM2x4x8x
AMD Instinct MI355X [VERIFY]288 GB
576
1152
2304
AMD Instinct MI325X256 GB
512
1024
2048
Apple Mac Studio (M3 Ultra, 256 GB)256 GB
512
1024
2048
AMD Instinct MI300X192 GB
384
768
1536
Apple Mac Studio (M2 Ultra, 192 GB)192 GB
384
768
1536
Apple Mac Pro (M2 Ultra, 192 GB)192 GB
384
768
1536

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull deepseek-v3:0324

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{deepseek-v3-0324-2025,
  title={DeepSeek V3-0324},
  author={DeepSeek AI},
  year={2025},
  url={https://huggingface.co/deepseek-ai/DeepSeek-V3-0324}
}
APA
DeepSeek AI (2025). DeepSeek V3-0324 [Model card]. Hugging Face. https://huggingface.co/deepseek-ai/DeepSeek-V3-0324
Plain text
DeepSeek V3-0324 (DeepSeek V3, DeepSeek AI, 2025). Available at https://huggingface.co/deepseek-ai/DeepSeek-V3-0324.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models