Llama 3.2Text LLMLlama 3.2 Community LicenseSep 2024

Llama 3.2 3B

Llama 3.2 3B punches well above its weight thanks to knowledge distillation from the 8B model. It's the sweet spot for AI PCs with NPUs (Intel Core Ultra, Snapdragon X, Ryzen AI) and runs at 50-100+ tok/s on midrange consumer GPUs. Tool-use and short-form summarization are particularly strong; it can struggle with long reasoning chains.

Parameters
3B
dense
Context
128K
tokens
Min VRAM (Q4)
1.8 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Llama 3.2 3B?

Llama 3.2 3B needs at minimum 1.8 GB of VRAM at Q4_K_M quantization (6.6 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A380 6GB (6 GB VRAM, $139 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Llama 3.2 3B (Llama 3.2, 3B params)As of 2024-09-25

TL;DR, what to buy

Recommended GPU
Intel Arc A380 6GB
6 GB VRAM · $139 MSRP
Min VRAM at Q4_K_M
1.8 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP166.6 GBReferenceTraining-precision weights
Q8_03.3 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K2.7 GBVery high6-bit, near Q8 quality
Q5_K_M2.3 GBHighStrong middle ground
Q4_K_M1.8 GBBalanced (recommended)Default for local deployments
Q4_01.9 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M1.4 GBLossyWhen VRAM is very tight
Q2_K1.1 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 6.6 GB; Q4_K_M needs 1.8 GB.

At full precision (FP16)

  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299
  • AMD RX 9060 XT 8GB8 GB · $299

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull llama3.2:3b
LM Studio
GGUF format
lms get bartowski/Llama-3.2-3B-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{llama-3-2-3b-2024,
  title={Llama 3.2 3B},
  author={Meta AI},
  year={2024},
  url={https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct}
}
APA
Meta AI (2024). Llama 3.2 3B [Model card]. Hugging Face. https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct
Plain text
Llama 3.2 3B (Llama 3.2, Meta AI, 2024). Available at https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models