Llama 3.1Text LLMLlama 3.1 Community LicenseJul 2024

Llama 3.1 8B

Llama 3.1 8B is Meta's best-in-class small dense model — competitive with much larger 13B-class models on reasoning and instruction following. It's fast enough for single-GPU local chat on 8 GB cards at Q4 and ships with full 128K context support. The weakness is that it still trails 70B-class models on multi-step reasoning and tool use.

Parameters
8B
dense
Context
128K
tokens
Min VRAM (Q4)
4.8 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Llama 3.1 8B?

Llama 3.1 8B needs at minimum 4.8 GB of VRAM at Q4_K_M quantization (17.6 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A580 8GB (8 GB VRAM, $179 MSRP). Measured throughput hits 1980 tok/s on 8× NVIDIA H100 SXM5 80GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Llama 3.1 8B (Llama 3.1, 8B params)As of 2024-07-23

TL;DR, what to buy

Recommended GPU
Intel Arc A580 8GB
8 GB VRAM · $179 MSRP
Min VRAM at Q4_K_M
4.8 GB
+ ~20-30% headroom for KV cache
Best measured speed
1980 tok/s
on 8× NVIDIA H100 SXM5 80GB

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1617.6 GBReferenceTraining-precision weights
Q8_08.8 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K7.1 GBVery high6-bit, near Q8 quality
Q5_K_M6.1 GBHighStrong middle ground
Q4_K_M4.8 GBBalanced (recommended)Default for local deployments
Q4_05.0 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M3.8 GBLossyWhen VRAM is very tight
Q2_K2.8 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 17.6 GB; Q4_K_M needs 4.8 GB.

At full precision (FP16)

  • Intel Arc Pro B6024 GB · $500
  • AMD RX 7900 XT20 GB · $749
  • Apple Mac mini (M4, 24 GB)24 GB · $799
  • AMD RX 7900 XTX24 GB · $899
  • NVIDIA RTX 309024 GB · $999
  • Apple Mac mini (M2, 24 GB)24 GB · $999
  • Apple Mac mini (M4, 32 GB)32 GB · $999
  • NVIDIA RTX 3090 Ti24 GB · $1,099
  • Apple MacBook Air 13" (M4, 24 GB)24 GB · $1,199
  • NVIDIA RTX 4000 Ada Generation20 GB · $1,250

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

85 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
8× NVIDIA H100 SXM5 80GB1980 tok/sFP1681925600W
Cerebras WSE-3 (CS-3)1850 tok/sFP16819223000W
Groq LPU (cloud)1850 tok/sFP168192350W
AMD Instinct MI300X 192GB920 tok/sQ4_K_M4096750W
Groq LPU (per chip)750 tok/sFP168192215W
NVIDIA GeForce RTX 5090 32GB720 tok/sQ4_K_M4096575W
NVIDIA GeForce RTX 4090 24GB540 tok/sQ4_K_M4096450W
NVIDIA B200 192GB520 tok/sFP1681921000W
Google TPU v5p410 tok/sFP168192700W
NVIDIA GH200 Grace Hopper 480GB380 tok/sFP1681921000W
NVIDIA B200 192GB360 tok/sQ4_K_M1310721000W
NVIDIA H200 141GB340 tok/sFP164096700W
AWS Trainium2320 tok/sFP168192500W
NVIDIA H100 SXM5 80GB282 tok/sFP164096700W
AMD Instinct MI300X 192GB260 tok/sFP168192750W
Google TPU v5e240 tok/sFP164096170W
NVIDIA A100 80GB SXM215 tok/sFP164096400W
AWS Trainium2215 tok/sQ4_K_M8192500W
NVIDIA H200 141GB215 tok/sQ4_K_M131072700W
NVIDIA A100 40GB195 tok/sFP164096400W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull llama3.1:8b
LM Studio
GGUF format
lms get meta-llama/Meta-Llama-3.1-8B-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{llama-3-1-8b-2024,
  title={Llama 3.1 8B},
  author={Meta AI},
  year={2024},
  url={https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct}
}
APA
Meta AI (2024). Llama 3.1 8B [Model card]. Hugging Face. https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct
Plain text
Llama 3.1 8B (Llama 3.1, Meta AI, 2024). Available at https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models