Llama 4Text LLMLlama 4 Community LicenseApr 2025

Llama 4 Maverick 17B

Llama 4 Maverick is Meta's flagship multimodal MoE — 400 B total parameters across 128 experts with 17 B activated per token. Benchmarks place it near GPT-4o on reasoning and beating Llama 3.1 405B at far lower per-token cost. Realistically it requires a 4x or 8x H100 / B200 cluster (~900 GB at FP16, ~245 GB at Q4_K_M). Not a homelab model.

Parameters
17B
17B active (MoE)
Context
1M
tokens
Min VRAM (Q4)
242 GB
weights only
Run locally?
NO
needs multi-GPU
Quick answer

What hardware do I need to run Llama 4 Maverick 17B?

Llama 4 Maverick 17B needs at minimum 242 GB of VRAM at Q4_K_M quantization (880 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Apple Mac Studio (M3 Ultra, 512 GB) (512 GB VRAM, $9,499 MSRP). Community benchmark submissions are open. This model exceeds 48 GB at Q4, so plan for a 2-or-more-GPU split.

Source: MyAIHardware model card: Llama 4 Maverick 17B (Llama 4, 17B params)As of 2025-04-05

TL;DR, what to buy

Recommended GPU
Apple Mac Studio (M3 Ultra, 512 GB)
512 GB VRAM · $9,499 MSRP
Min VRAM at Q4_K_M
242 GB
+ ~20-30% headroom for KV cache
Best measured speed
no community benchmarks yet
submit yours below

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP16880.0 GBReferenceTraining-precision weights
Q8_0440.0 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K356.4 GBVery high6-bit, near Q8 quality
Q5_K_M303.6 GBHighStrong middle ground
Q4_K_M242.0 GBBalanced (recommended)Default for local deployments
Q4_0247.5 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M189.2 GBLossyWhen VRAM is very tight
Q2_K140.8 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 880.0 GB; Q4_K_M needs 242.0 GB.

At full precision (FP16)

No GPU in our database has enough VRAM for FP16. Use Q4_K_M or multi-GPU.

At Q4_K_M quantization

  • Apple Mac Studio (M3 Ultra, 256 GB)256 GB · $5,599
  • Apple Mac Studio (M3 Ultra, 512 GB)512 GB · $9,499
  • AMD Instinct MI325X256 GB · $18,000
  • AMD Instinct MI355X [VERIFY]288 GB · $25,000
  • NVIDIA GH200 Grace Hopper576 GB · $65,000

Multi-GPU splits (Q4_K_M target: 242 GB)

GPUPer-card VRAM2x4x8x
AMD Instinct MI300X192 GB
384
768
1536
Apple Mac Studio (M2 Ultra, 192 GB)192 GB
384
768
1536
Apple Mac Pro (M2 Ultra, 192 GB)192 GB
384
768
1536
NVIDIA B200180 GB
360
720
1440
NVIDIA B100180 GB
360
720
1440
NVIDIA H200141 GB
282
564
1128

Community benchmarks

We don't have community benchmarks for this exact model yet.

No benchmarks yet for this exact model. Submit your own measurement.

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull llama4:maverick

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{llama-4-maverick-2025,
  title={Llama 4 Maverick 17B},
  author={Meta AI},
  year={2025},
  url={https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct}
}
APA
Meta AI (2025). Llama 4 Maverick 17B [Model card]. Hugging Face. https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct
Plain text
Llama 4 Maverick 17B (Llama 4, Meta AI, 2025). Available at https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models