Llama 3.1Text LLMLlama 3.1 Community LicenseJul 2024

Llama 3.1 70B

Llama 3.1 70B is the workhorse of self-hosted AI in 2024-2026 — competitive with GPT-4-class proprietary models on most benchmarks. At Q4_K_M it fits in a single RTX 4090 or 5090 (24-32 GB) with reduced context, or comfortably on dual 24 GB cards. The trade-off versus 8B is roughly 5-8x slower tok/s on identical hardware.

Parameters
70B
dense
Context
128K
tokens
Min VRAM (Q4)
42.4 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Llama 3.1 70B?

Llama 3.1 70B needs at minimum 42.4 GB of VRAM at Q4_K_M quantization (154 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac mini (M4 Pro, 64 GB) (64 GB VRAM, $2,199 MSRP). Measured throughput hits 540 tok/s on NVIDIA H200 141GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Llama 3.1 70B (Llama 3.1, 70B params)As of 2024-07-23

TL;DR, what to buy

Recommended GPU
Apple Mac mini (M4 Pro, 64 GB)
64 GB VRAM · $2,199 MSRP
Min VRAM at Q4_K_M
42.4 GB
+ ~20-30% headroom for KV cache
Best measured speed
540 tok/s
on NVIDIA H200 141GB

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP16154.0 GBReferenceTraining-precision weights
Q8_077.0 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K62.4 GBVery high6-bit, near Q8 quality
Q5_K_M53.1 GBHighStrong middle ground
Q4_K_M42.4 GBBalanced (recommended)Default for local deployments
Q4_043.3 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M33.1 GBLossyWhen VRAM is very tight
Q2_K24.6 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 154.0 GB; Q4_K_M needs 42.4 GB.

At full precision (FP16)

  • Apple Mac Studio (M3 Ultra, 256 GB)256 GB · $5,599
  • Apple Mac Studio (M2 Ultra, 192 GB)192 GB · $6,599
  • Apple Mac Studio (M3 Ultra, 512 GB)512 GB · $9,499
  • Apple Mac Pro (M2 Ultra, 192 GB)192 GB · $9,599
  • AMD Instinct MI300X192 GB · $14,999
  • AMD Instinct MI325X256 GB · $18,000
  • AMD Instinct MI355X [VERIFY]288 GB · $25,000
  • NVIDIA B100180 GB · $38,000
  • NVIDIA B200180 GB · $40,000
  • NVIDIA GH200 Grace Hopper576 GB · $65,000

At Q4_K_M quantization

  • Apple Mac mini (M4 Pro, 48 GB)48 GB · $1,799
  • Apple Mac mini (M4 Pro, 64 GB)64 GB · $2,199
  • Apple Mac Studio (M1 Max, 64 GB)64 GB · $2,399
  • Apple Mac Studio (M2 Max, 64 GB)64 GB · $2,399
  • Apple MacBook Pro 14" (M4 Pro, 48 GB)48 GB · $2,399
  • Apple Mac Studio (M4 Max, 64 GB)64 GB · $2,499
  • Apple MacBook Pro 16" (M4 Pro, 48 GB)48 GB · $2,499
  • Intel Data Center GPU Max 1550128 GB · $2,500
  • Apple Mac Studio (M2 Max, 96 GB)96 GB · $2,999
  • Apple Mac Studio (M4 Max, 128 GB)128 GB · $3,499
  • Apple MacBook Pro 16" (M1 Max, 64 GB)64 GB · $3,499
  • Apple MacBook Pro 16" (M3 Max, 64 GB)64 GB · $3,499

Community benchmarks

52 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
NVIDIA H200 141GB540 tok/sQ4_K_M4096700W
Cerebras WSE-3 (CS-3)450 tok/sFP16819223000W
8× NVIDIA H100 SXM5 80GB412 tok/sQ4_K_M81925600W
NVIDIA H100 SXM5 80GB410 tok/sQ4_K_M4096700W
AMD Instinct MI300X 192GB380 tok/sQ4_K_M4096750W
Groq LPU (8-chip rack)285 tok/sFP1681921720W
NVIDIA B200 192GB138 tok/sQ4_K_M81921000W
Google TPU v5p110 tok/sINT88192700W
NVIDIA B200 192GB98 tok/sQ8_081921000W
NVIDIA GH200 Grace Hopper 480GB95 tok/sQ4_K_M81921000W
AWS Trainium292 tok/sINT88192500W
NVIDIA H200 141GB84 tok/sQ4_K_M8192700W
NVIDIA GeForce RTX 5090 32GB78 tok/sQ4_K_M4096575W
AMD Instinct MI300X 192GB72 tok/sQ4_K_M8192750W
NVIDIA H100 SXM5 80GB66 tok/sQ4_K_M4096700W
Google TPU v5e62 tok/sINT84096170W
NVIDIA H200 141GB58 tok/sQ8_08192700W
AWS Trainium258 tok/sQ4_K_M8192500W
AMD Instinct MI250X 128GB52 tok/sQ4_K_M8192560W
AMD Instinct MI300X 192GB48 tok/sQ8_08192750W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull llama3.1:70b
LM Studio
GGUF format
lms get bartowski/Meta-Llama-3.1-70B-Instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{llama-3-1-70b-2024,
  title={Llama 3.1 70B},
  author={Meta AI},
  year={2024},
  url={https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct}
}
APA
Meta AI (2024). Llama 3.1 70B [Model card]. Hugging Face. https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct
Plain text
Llama 3.1 70B (Llama 3.1, Meta AI, 2024). Available at https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models