GraniteText LLMApache 2.0Feb 2025

Granite 3.2 8B Instruct

Granite 3.2 8B Instruct is IBM's February 2025 reasoning model — same 8 B footprint as Llama 3.1 8B but trained on a fully audited corpus that IBM will indemnify enterprise customers against. It ships with a toggleable thinking mode for chain-of-thought reasoning and matches or beats Llama 3.1 8B on most leaderboards. The pitch is not raw quality, it is the legal and procurement story: Apache 2.0 weights, documented training data, and IBM's enterprise SLA on watsonx.ai for buyers who cannot deploy Meta or Alibaba weights. Runs at 60-100 tok/s on a single RTX 4060 8 GB at Q4_K_M. For a regulated bank in Istanbul or a public-sector buyer in Lagos, this is often the only 8B model that clears procurement.

Parameters
8B
dense
Context
128K
tokens
Min VRAM (Q4)
4.8 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Granite 3.2 8B Instruct?

Granite 3.2 8B Instruct needs at minimum 4.8 GB of VRAM at Q4_K_M quantization (17.6 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A580 8GB (8 GB VRAM, $179 MSRP). Community benchmark submissions are open. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Granite 3.2 8B Instruct (Granite, 8B params)As of 2025-02-26

TL;DR, what to buy

Recommended GPU
Intel Arc A580 8GB
8 GB VRAM · $179 MSRP
Min VRAM at Q4_K_M
4.8 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1617.6 GBReferenceTraining-precision weights
Q8_08.8 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K7.1 GBVery high6-bit, near Q8 quality
Q5_K_M6.1 GBHighStrong middle ground
Q4_K_M4.8 GBBalanced (recommended)Default for local deployments
Q4_05.0 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M3.8 GBLossyWhen VRAM is very tight
Q2_K2.8 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 17.6 GB; Q4_K_M needs 4.8 GB.

At full precision (FP16)

  • Intel Arc Pro B6024 GB · $500
  • AMD RX 7900 XT20 GB · $749
  • Apple Mac mini (M4, 24 GB)24 GB · $799
  • AMD RX 7900 XTX24 GB · $899
  • NVIDIA RTX 309024 GB · $999
  • Apple Mac mini (M2, 24 GB)24 GB · $999
  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • NVIDIA RTX 3090 Ti24 GB · $1,099
  • Apple MacBook Air 13" (M4, 24 GB)24 GB · $1,199
  • NVIDIA RTX 4000 Ada Generation20 GB · $1,250

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull granite3.2
LM Studio
GGUF format
lms get ibm-granite/granite-3.2-8b-instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{granite-3-2-8b-2025,
  title={Granite 3.2 8B Instruct},
  author={Granite},
  year={2025},
  url={https://huggingface.co/ibm-granite/granite-3.2-8b-instruct}
}
APA
Granite (2025). Granite 3.2 8B Instruct [Model card]. Hugging Face. https://huggingface.co/ibm-granite/granite-3.2-8b-instruct
Plain text
Granite 3.2 8B Instruct (Granite, Granite, 2025). Available at https://huggingface.co/ibm-granite/granite-3.2-8b-instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models