Llama 3.2Vision LLMLlama 3.2 Community LicenseSep 2024

Llama 3.2 Vision 11B

Llama 3.2 Vision 11B is Meta's accessible open multimodal model — handles image-grounded QA, chart reading, and document understanding well. Fits in 16 GB VRAM at Q4_K_M. Competitive with GPT-4V on common tasks but lags on dense-text OCR and very detailed visual reasoning where the 90B variant pulls ahead.

Parameters
11B
dense
Context
128K
tokens
Min VRAM (Q4)
6.7 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Llama 3.2 Vision 11B?

Llama 3.2 Vision 11B needs at minimum 6.7 GB of VRAM at Q4_K_M quantization (24.2 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc B570 (10 GB VRAM, $219 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Llama 3.2 Vision 11B (Llama 3.2, 11B params)As of 2024-09-25

TL;DR, what to buy

Recommended GPU
Intel Arc B570
10 GB VRAM · $219 MSRP
Min VRAM at Q4_K_M
6.7 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1624.2 GBReferenceTraining-precision weights
Q8_012.1 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K9.8 GBVery high6-bit, near Q8 quality
Q5_K_M8.3 GBHighStrong middle ground
Q4_K_M6.7 GBBalanced (recommended)Default for local deployments
Q4_06.8 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M5.2 GBLossyWhen VRAM is very tight
Q2_K3.9 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 24.2 GB; Q4_K_M needs 6.7 GB.

At full precision (FP16)

  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • AMD Radeon AI PRO R970032 GB · $1,299
  • Apple MacBook Air 13" (M4, 32 GB)32 GB · $1,399
  • Apple Mac mini (M2 Pro, 32 GB)32 GB · $1,699
  • Apple Mac mini (M4 Pro, 48 GB)48 GB · $1,799
  • NVIDIA RTX 509032 GB · $1,999
  • Apple Mac Studio (M1 Max, 32 GB)32 GB · $1,999
  • Apple Mac Studio (M2 Max, 32 GB)32 GB · $1,999
  • Apple Mac Studio (M4 Max, 36 GB)36 GB · $1,999
  • Apple Mac mini (M4 Pro, 64 GB)64 GB · $2,199

At Q4_K_M quantization

  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299
  • AMD RX 9060 XT 8GB8 GB · $299
  • NVIDIA GeForce RTX 2060 (12 GB refresh)12 GB · $299
  • Intel Arc Pro B5016 GB · $299

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull llama3.2-vision:11b
LM Studio
GGUF format
lms get bartowski/Llama-3.2-11B-Vision-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{llama-3-2-vision-11b-2024,
  title={Llama 3.2 Vision 11B},
  author={Meta AI},
  year={2024},
  url={https://huggingface.co/meta-llama/Llama-3.2-11B-Vision-Instruct}
}
APA
Meta AI (2024). Llama 3.2 Vision 11B [Model card]. Hugging Face. https://huggingface.co/meta-llama/Llama-3.2-11B-Vision-Instruct
Plain text
Llama 3.2 Vision 11B (Llama 3.2, Meta AI, 2024). Available at https://huggingface.co/meta-llama/Llama-3.2-11B-Vision-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models