Qwen 2.5Text LLMApache 2.0Sep 2024

Qwen 2.5 14B

Qwen 2.5 14B sits in the 'serious assistant' sweet spot — clearly better than 7B on multi-step reasoning while still fitting comfortably in 16 GB VRAM at Q4_K_M. It's a good choice for a single RTX 4070 Ti Super or 4080 build. Trade-off vs 7B is roughly 2x slower tok/s.

Parameters
14B
dense
Context
128K
tokens
Min VRAM (Q4)
8.5 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Qwen 2.5 14B?

Qwen 2.5 14B needs at minimum 8.5 GB of VRAM at Q4_K_M quantization (30.8 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A580 12GB (variant) [VERIFY] (12 GB VRAM, $219 MSRP). Measured throughput hits 1080 tok/s on 8× NVIDIA H100 SXM5 80GB (DGX H100). This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Qwen 2.5 14B (Qwen 2.5, 14B params)As of 2024-09-19

TL;DR, what to buy

Recommended GPU
Intel Arc A580 12GB (variant) [VERIFY]
12 GB VRAM · $219 MSRP
Min VRAM at Q4_K_M
8.5 GB
+ ~20-30% headroom for KV cache
Best measured speed
1080 tok/s
on 8× NVIDIA H100 SXM5 80GB (DGX H100)

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1630.8 GBReferenceTraining-precision weights
Q8_015.4 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K12.5 GBVery high6-bit, near Q8 quality
Q5_K_M10.6 GBHighStrong middle ground
Q4_K_M8.5 GBBalanced (recommended)Default for local deployments
Q4_08.7 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M6.6 GBLossyWhen VRAM is very tight
Q2_K4.9 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 30.8 GB; Q4_K_M needs 8.5 GB.

At full precision (FP16)

  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • AMD Radeon AI PRO R970032 GB · $1,299
  • Apple MacBook Air 13" (M4, 32 GB)32 GB · $1,399
  • Apple Mac mini (M2 Pro, 32 GB)32 GB · $1,699
  • Apple Mac mini (M4 Pro, 48 GB)48 GB · $1,799
  • NVIDIA RTX 509032 GB · $1,999
  • Apple Mac Studio (M1 Max, 32 GB)32 GB · $1,999
  • Apple Mac Studio (M2 Max, 32 GB)32 GB · $1,999
  • Apple Mac Studio (M4 Max, 36 GB)36 GB · $1,999
  • Apple Mac mini (M4 Pro, 64 GB)64 GB · $2,199

At Q4_K_M quantization

  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 2060 (12 GB refresh)12 GB · $299
  • Intel Arc Pro B5016 GB · $299
  • AMD RX 7600 XT16 GB · $329
  • Intel Arc A770 16GB16 GB · $329
  • NVIDIA GeForce RTX 3060 (12 GB)12 GB · $329
  • AMD RX 9060 XT 16GB16 GB · $349
  • Intel Arc B770 16GB [VERIFY]16 GB · $349
  • AMD RX 7700 XT12 GB · $399
  • NVIDIA RTX 5060 Ti 16GB16 GB · $429

Community benchmarks

44 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
8× NVIDIA H100 SXM5 80GB (DGX H100)1080 tok/sQ4_K_M81925600W
Groq LPU (8-chip rack)480 tok/sFP1681921720W
NVIDIA B200 192GB285 tok/sQ4_K_M81921000W
NVIDIA B200 192GB268 tok/sQ4_K_M81921000W
Google TPU v5p (Trillium)195 tok/sINT88192300W
NVIDIA H200 141GB178 tok/sQ4_K_M8192700W
NVIDIA H100 SXM5 80GB168 tok/sQ4_K_M8192700W
NVIDIA H100 SXM5 80GB165 tok/sQ4_K_M8192700W
AWS Trainium2 Trn2158 tok/sINT88192350W
NVIDIA H100 SXM5 80GB145 tok/sQ4_K_M8192700W
AMD Instinct MI300X 192GB142 tok/sQ4_K_M8192750W
AMD Instinct MI300X 192GB142 tok/sQ4_K_M8192750W
AMD Instinct MI300X 192GB128 tok/sQ4_K_M8192750W
NVIDIA GeForce RTX 5090 32GB115 tok/sQ4_K_M8192575W
NVIDIA A100 SXM4 80GB108 tok/sQ4_K_M8192400W
NVIDIA GeForce RTX 5090 32GB105 tok/sQ4_K_M8192575W
NVIDIA GeForce RTX 5090 32GB92 tok/sQ4_K_M8192575W
NVIDIA L40S 48GB92 tok/sQ4_K_M8192350W
2× NVIDIA RTX 4090 24GB92 tok/sQ4_K_M8192900W
NVIDIA GeForce RTX 5080 16GB86 tok/sQ4_K_M8192360W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull qwen2.5:14b
LM Studio
GGUF format
lms get Qwen/Qwen2.5-14B-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{qwen-2-5-14b-2024,
  title={Qwen 2.5 14B},
  author={Alibaba Cloud Qwen Team},
  year={2024},
  url={https://huggingface.co/Qwen/Qwen2.5-14B-Instruct}
}
APA
Alibaba Cloud Qwen Team (2024). Qwen 2.5 14B [Model card]. Hugging Face. https://huggingface.co/Qwen/Qwen2.5-14B-Instruct
Plain text
Qwen 2.5 14B (Qwen 2.5, Alibaba Cloud Qwen Team, 2024). Available at https://huggingface.co/Qwen/Qwen2.5-14B-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models