DeepSeek CoderCode LLMDeepSeek LicenseJun 2024

DeepSeek-Coder-V2 16B

DeepSeek-Coder-V2 16B (Lite) is a 16 B MoE with 2.4 B activated params, designed for fast code completion. It matches GPT-4 Turbo on HumanEval and supports 338 programming languages with a 128K context window. At Q4_K_M it runs at 60-100 tok/s on a single RTX 4090, making it the strongest local code assistant in this class. The bigger 236 B Coder-V2 variant exists for datacenter use.

Parameters
16B
2.4B active (MoE)
Context
160K
tokens
Min VRAM (Q4)
9.7 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run DeepSeek-Coder-V2 16B?

DeepSeek-Coder-V2 16B needs at minimum 9.7 GB of VRAM at Q4_K_M quantization (35.2 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc Pro B50 (16 GB VRAM, $299 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: DeepSeek-Coder-V2 16B (DeepSeek Coder, 16B params)As of 2024-06-17

TL;DR, what to buy

Recommended GPU
Intel Arc Pro B50
16 GB VRAM · $299 MSRP
Min VRAM at Q4_K_M
9.7 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1635.2 GBReferenceTraining-precision weights
Q8_017.6 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K14.3 GBVery high6-bit, near Q8 quality
Q5_K_M12.1 GBHighStrong middle ground
Q4_K_M9.7 GBBalanced (recommended)Default for local deployments
Q4_09.9 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M7.6 GBLossyWhen VRAM is very tight
Q2_K5.6 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 35.2 GB; Q4_K_M needs 9.7 GB.

At full precision (FP16)

  • Apple Mac mini (M4 Pro, 48 GB)48 GB · $1,799
  • Apple Mac Studio (M4 Max, 36 GB)36 GB · $1,999
  • Apple Mac mini (M4 Pro, 64 GB)64 GB · $2,199
  • Apple Mac Studio (M1 Max, 64 GB)64 GB · $2,399
  • Apple Mac Studio (M2 Max, 64 GB)64 GB · $2,399
  • Apple MacBook Pro 14" (M3 Pro, 36 GB)36 GB · $2,399
  • Apple MacBook Pro 14" (M4 Pro, 48 GB)48 GB · $2,399
  • Apple Mac Studio (M4 Max, 64 GB)64 GB · $2,499
  • Apple MacBook Pro 16" (M4 Pro, 48 GB)48 GB · $2,499
  • Intel Data Center GPU Max 1550128 GB · $2,500

At Q4_K_M quantization

  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 2060 (12 GB refresh)12 GB · $299
  • Intel Arc Pro B5016 GB · $299
  • AMD RX 7600 XT16 GB · $329
  • Intel Arc A770 16GB16 GB · $329
  • NVIDIA GeForce RTX 3060 (12 GB)12 GB · $329
  • AMD RX 9060 XT 16GB16 GB · $349
  • Intel Arc B770 16GB [VERIFY]16 GB · $349
  • AMD RX 7700 XT12 GB · $399
  • NVIDIA RTX 5060 Ti 16GB16 GB · $429

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull deepseek-coder-v2
LM Studio
GGUF format
lms get bartowski/DeepSeek-Coder-V2-Lite-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{deepseek-coder-v2-16b-2024,
  title={DeepSeek-Coder-V2 16B},
  author={DeepSeek AI},
  year={2024},
  url={https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct}
}
APA
DeepSeek AI (2024). DeepSeek-Coder-V2 16B [Model card]. Hugging Face. https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
Plain text
DeepSeek-Coder-V2 16B (DeepSeek Coder, DeepSeek AI, 2024). Available at https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models