Interactive tool

LLM VRAM Calculator , Will it fit on your GPU?

Pick a model, a quantization, and a context length. We'll tell you exactly how much VRAM you need, which GPUs in our database can run it, and what to do when nothing fits.

30 models10 quantizations189 GPUs in compatibility list

Configure

Set your workload, see the numbers update live.

Meta's flagship 70B model. Competitive with GPT-4 on many tasks.

Q4_K_M — 4-bit medium, popular

8K tokens
2K4K8K16K32K64K128K
1
18162432

You need approximately

0.0 GB

of VRAM to run Llama 3 70B at Q4_K_M.

8K context
1 concurrent user
FP16 KV cache
Model weights
42.4 GB
KV cache
20 GB
Activations + overhead
2.1 GB

GPU compatibility

31 of 189 GPUs in our database can run this.

31 fits158 miss
NVIDIA H100 SXM5
NVIDIA · datacenter
80 GB
+15.5 GB
NVIDIA A100 80GB
NVIDIA · datacenter
80 GB
+15.5 GB
NVIDIA H100 NVL
NVIDIA · datacenter
94 GB
+29.5 GB
Apple Mac Studio (M2 Max, 96 GB)
Apple · apple-silicon
96 GB
+31.5 GB
Apple Mac Studio (M3 Ultra, 96 GB)
Apple · apple-silicon
96 GB
+31.5 GB
NVIDIA RTX PRO 6000 Blackwell
NVIDIA · workstation
96 GB
+31.5 GB
Apple MacBook Pro 16" (M2 Max, 96 GB)
Apple · apple-silicon
96 GB
+31.5 GB
NVIDIA H20 (China export SKU) [VERIFY]
NVIDIA · datacenter
96 GB
+31.5 GB
Intel Gaudi 3 PCIe [VERIFY]
Intel · datacenter
96 GB
+31.5 GB
NVIDIA RTX PRO 6000 Blackwell Max-Q
NVIDIA · workstation
96 GB
+31.5 GB
AMD Instinct MI300A
AMD · datacenter
128 GB
+63.5 GB
AMD Instinct MI250X
AMD · datacenter
128 GB
+63.5 GB
Intel Data Center GPU Max 1550
Intel · datacenter
128 GB
+63.5 GB
Apple Mac Studio (M1 Ultra, 128 GB)
Apple · apple-silicon
128 GB
+63.5 GB
Apple Mac Studio (M2 Ultra, 128 GB)
Apple · apple-silicon
128 GB
+63.5 GB
Apple Mac Studio (M4 Max, 128 GB)
Apple · apple-silicon
128 GB
+63.5 GB
Apple MacBook Pro 16" (M3 Max, 128 GB)
Apple · apple-silicon
128 GB
+63.5 GB
Apple MacBook Pro 16" (M4 Max, 128 GB)
Apple · apple-silicon
128 GB
+63.5 GB
AMD Instinct MI250
AMD · datacenter
128 GB
+63.5 GB
NVIDIA H200
NVIDIA · datacenter
141 GB
+76.5 GB
NVIDIA H200 NVL
NVIDIA · datacenter
141 GB
+76.5 GB
NVIDIA B200
NVIDIA · datacenter
180 GB
+115.5 GB
NVIDIA B100
NVIDIA · datacenter
180 GB
+115.5 GB
AMD Instinct MI300X
AMD · datacenter
192 GB
+127.5 GB
Apple Mac Studio (M2 Ultra, 192 GB)
Apple · apple-silicon
192 GB
+127.5 GB
Apple Mac Pro (M2 Ultra, 192 GB)
Apple · apple-silicon
192 GB
+127.5 GB
AMD Instinct MI325X
AMD · datacenter
256 GB
+191.5 GB
Apple Mac Studio (M3 Ultra, 256 GB)
Apple · apple-silicon
256 GB
+191.5 GB
AMD Instinct MI355X [VERIFY]
AMD · datacenter
288 GB
+223.5 GB
Apple Mac Studio (M3 Ultra, 512 GB)
Apple · apple-silicon
512 GB
+447.5 GB
NVIDIA GH200 Grace Hopper
NVIDIA · datacenter
576 GB
+511.5 GB
Apple Mac mini (M4 Pro, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Studio (M1 Max, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Studio (M1 Ultra, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Studio (M2 Max, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Studio (M2 Ultra, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Studio (M4 Max, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple MacBook Pro 16" (M1 Max, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple MacBook Pro 16" (M3 Max, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
Apple Mac Pro (M2 Ultra, 64 GB)
Apple · apple-silicon
64 GB
-0.5 GB
AMD Instinct MI210
AMD · datacenter
64 GB
-0.5 GB
NVIDIA L40S
NVIDIA · datacenter
48 GB
-16.5 GB
NVIDIA RTX 6000 Ada
NVIDIA · workstation
48 GB
-16.5 GB
NVIDIA RTX A6000
NVIDIA · workstation
48 GB
-16.5 GB
Apple Mac mini (M4 Pro, 48 GB)
Apple · apple-silicon
48 GB
-16.5 GB
NVIDIA RTX PRO 5000 Blackwell
NVIDIA · workstation
48 GB
-16.5 GB
Apple MacBook Pro 14" (M4 Pro, 48 GB)
Apple · apple-silicon
48 GB
-16.5 GB
Apple MacBook Pro 16" (M4 Pro, 48 GB)
Apple · apple-silicon
48 GB
-16.5 GB
Apple MacBook Pro 16" (M4 Max, 48 GB)
Apple · apple-silicon
48 GB
-16.5 GB
NVIDIA B40 PCIe [VERIFY]
NVIDIA · datacenter
48 GB
-16.5 GB
AMD Radeon PRO W7900
AMD · workstation
48 GB
-16.5 GB
NVIDIA A100 40GB
NVIDIA · datacenter
40 GB
-24.5 GB
Apple Mac Studio (M4 Max, 36 GB)
Apple · apple-silicon
36 GB
-28.5 GB
Apple MacBook Pro 14" (M3 Pro, 36 GB)
Apple · apple-silicon
36 GB
-28.5 GB
Apple MacBook Pro 16" (M3 Max, 36 GB)
Apple · apple-silicon
36 GB
-28.5 GB
Apple MacBook Pro 16" (M4 Max, 36 GB)
Apple · apple-silicon
36 GB
-28.5 GB
NVIDIA RTX 5090
NVIDIA · consumer
32 GB
-32.5 GB
Apple Mac mini (M2 Pro, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple Mac mini (M4, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple Mac Studio (M1 Max, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple Mac Studio (M2 Max, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
NVIDIA RTX PRO 4500 Blackwell
NVIDIA · workstation
32 GB
-32.5 GB
NVIDIA RTX 5000 Ada Generation
NVIDIA · workstation
32 GB
-32.5 GB
Apple MacBook Air 13" (M4, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple MacBook Pro 14" (M1 Pro, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple MacBook Pro 16" (M1 Max, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple MacBook Pro 14" (M2 Pro, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
Apple MacBook Pro 16" (M2 Max, 32 GB)
Apple · apple-silicon
32 GB
-32.5 GB
AMD Radeon PRO W7800
AMD · workstation
32 GB
-32.5 GB
AMD Radeon PRO W6800
AMD · workstation
32 GB
-32.5 GB
AMD Radeon AI PRO R9700
AMD · workstation
32 GB
-32.5 GB
NVIDIA RTX A5500
NVIDIA · workstation
24 GB
-40.5 GB
NVIDIA RTX 4090
NVIDIA · consumer
24 GB
-40.5 GB
NVIDIA RTX 3090 Ti
NVIDIA · consumer
24 GB
-40.5 GB
NVIDIA RTX 3090
NVIDIA · consumer
24 GB
-40.5 GB
AMD RX 7900 XTX
AMD · consumer
24 GB
-40.5 GB
Apple Mac mini (M2, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple Mac mini (M4, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple Mac mini (M4 Pro, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
NVIDIA RTX PRO 4000 Blackwell
NVIDIA · workstation
24 GB
-40.5 GB
NVIDIA RTX 4500 Ada Generation
NVIDIA · workstation
24 GB
-40.5 GB
NVIDIA GeForce RTX 5090 Laptop GPU
NVIDIA · consumer
24 GB
-40.5 GB
Apple MacBook Air 13" (M2, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple MacBook Air 13" (M3, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple MacBook Air 13" (M4, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple MacBook Pro 14" (M4 Pro, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
Apple iMac 24" (M4, 24 GB)
Apple · apple-silicon
24 GB
-40.5 GB
NVIDIA TITAN RTX
NVIDIA · workstation
24 GB
-40.5 GB
Intel Arc Pro B60
Intel · workstation
24 GB
-40.5 GB
AMD RX 7900 XT
AMD · consumer
20 GB
-44.5 GB
NVIDIA RTX 4000 Ada Generation
NVIDIA · workstation
20 GB
-44.5 GB
Apple MacBook Pro 14" (M3 Pro, 18 GB)
Apple · apple-silicon
18 GB
-46.5 GB
NVIDIA RTX A5000
NVIDIA · workstation
16 GB
-48.5 GB
NVIDIA RTX 5080
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 5070 Ti
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 5060 Ti 16GB
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 4080 Super
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 4080
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 4070 Ti Super
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX 4060 Ti 16GB
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA GeForce RTX 3080 Laptop GPU (16 GB)
NVIDIA · consumer
16 GB
-48.5 GB
AMD RX 7800 XT
AMD · consumer
16 GB
-48.5 GB
AMD RX 9070 XT
AMD · consumer
16 GB
-48.5 GB
AMD RX 9070
AMD · consumer
16 GB
-48.5 GB
AMD RX 9060 XT 16GB
AMD · consumer
16 GB
-48.5 GB
AMD RX 7600 XT
AMD · consumer
16 GB
-48.5 GB
Intel Arc A770 16GB
Intel · consumer
16 GB
-48.5 GB
Apple Mac mini (M1, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple Mac mini (M2, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple Mac mini (M2 Pro, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple Mac mini (M4, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
NVIDIA GeForce RTX 5080 Laptop GPU
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA GeForce RTX 4090 Laptop GPU
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA GeForce RTX 3080 Ti Laptop GPU (16 GB)
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA GeForce RTX 3080 Mobile (16 GB sibling — "3090M" OEM badge) [VERIFY]
NVIDIA · consumer
16 GB
-48.5 GB
NVIDIA RTX A5500 Laptop GPU
NVIDIA · workstation
16 GB
-48.5 GB
NVIDIA RTX A5000 Laptop GPU
NVIDIA · workstation
16 GB
-48.5 GB
Apple MacBook Air 13" (M1, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple MacBook Air 13" (M2, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple MacBook Air 13" (M3, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple MacBook Air 13" (M4, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple MacBook Pro 14" (M1 Pro, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple MacBook Pro 14" (M2 Pro, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple iMac 24" (M1, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple iMac 24" (M3, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple iMac 24" (M4, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple iPad Pro (M4, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
Apple Vision Pro (M2, 16 GB)
Apple · apple-silicon
16 GB
-48.5 GB
AMD Radeon RX 7900 GRE
AMD · consumer
16 GB
-48.5 GB
AMD Radeon RX 6950 XT
AMD · consumer
16 GB
-48.5 GB
AMD Radeon RX 6900 XT
AMD · consumer
16 GB
-48.5 GB
AMD Radeon RX 6800 XT
AMD · consumer
16 GB
-48.5 GB
AMD Radeon RX 6800
AMD · consumer
16 GB
-48.5 GB
Intel Arc B770 16GB [VERIFY]
Intel · consumer
16 GB
-48.5 GB
Intel Data Center GPU Flex 170
Intel · datacenter
16 GB
-48.5 GB
Intel Arc Pro B50
Intel · workstation
16 GB
-48.5 GB
NVIDIA RTX 5070
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA RTX 4070 Ti
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA RTX 4070 Super
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA RTX 4070
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA RTX 3080 Ti
NVIDIA · consumer
12 GB
-52.5 GB
AMD RX 7700 XT
AMD · consumer
12 GB
-52.5 GB
Intel Arc B580
Intel · consumer
12 GB
-52.5 GB
NVIDIA GeForce RTX 5070 Ti Laptop GPU
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA GeForce RTX 4080 Laptop GPU
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA GeForce RTX 2060 (12 GB refresh)
NVIDIA · consumer
12 GB
-52.5 GB
NVIDIA RTX A3000 Laptop GPU (12 GB)
NVIDIA · workstation
12 GB
-52.5 GB
NVIDIA GeForce RTX 3060 (12 GB)
NVIDIA · consumer
12 GB
-52.5 GB
AMD Radeon RX 6700 XT
AMD · consumer
12 GB
-52.5 GB
Intel Arc A580 12GB (variant) [VERIFY]
Intel · consumer
12 GB
-52.5 GB
NVIDIA GeForce RTX 2080 Ti
NVIDIA · consumer
11 GB
-53.5 GB
NVIDIA GeForce GTX 1080 Ti
NVIDIA · consumer
11 GB
-53.5 GB
Intel Arc B570
Intel · consumer
10 GB
-54.5 GB
NVIDIA RTX 5060 Ti 8GB
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA RTX 5060 8GB
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA RTX 4060 Ti 8GB
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA RTX 4060 8GB
NVIDIA · consumer
8 GB
-56.5 GB
AMD RX 9060 XT 8GB
AMD · consumer
8 GB
-56.5 GB
AMD RX 7600
AMD · consumer
8 GB
-56.5 GB
Intel Arc A750 8GB
Intel · consumer
8 GB
-56.5 GB
Intel Arc A580 8GB
Intel · consumer
8 GB
-56.5 GB
Apple Mac mini (M1, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
Apple Mac mini (M2, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
NVIDIA GeForce RTX 5070 Laptop GPU
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 4070 Laptop GPU
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 4060 Laptop GPU
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 2070 SUPER
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 3080 Laptop GPU (8 GB)
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA RTX A4500 Laptop GPU
NVIDIA · workstation
8 GB
-56.5 GB
Apple MacBook Air 13" (M1, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
Apple MacBook Air 13" (M2, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
Apple MacBook Air 13" (M3, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
NVIDIA RTX A2000 Laptop GPU (8 GB)
NVIDIA · workstation
8 GB
-56.5 GB
Apple iMac 24" (M1, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
Apple iMac 24" (M3, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
Apple iPad Pro (M4, 8 GB)
Apple · apple-silicon
8 GB
-56.5 GB
NVIDIA GeForce RTX 3070 Ti
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 3070
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 3060 Ti
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 3050 (8 GB)
NVIDIA · consumer
8 GB
-56.5 GB
NVIDIA GeForce RTX 2080 SUPER
NVIDIA · consumer
8 GB
-56.5 GB
AMD Radeon PRO W6600
AMD · workstation
8 GB
-56.5 GB
AMD Radeon PRO W5700
AMD · workstation
8 GB
-56.5 GB
NVIDIA GeForce RTX 3060 Laptop GPU
NVIDIA · consumer
6 GB
-58.5 GB
NVIDIA GeForce GTX 1660 SUPER
NVIDIA · consumer
6 GB
-58.5 GB
NVIDIA GeForce RTX 3050 (6 GB)
NVIDIA · consumer
6 GB
-58.5 GB
Intel Arc A380 6GB
Intel · consumer
6 GB
-58.5 GB
NVIDIA GeForce RTX 3050 Laptop GPU
NVIDIA · consumer
4 GB
-60.5 GB
NVIDIA RTX A1000 Laptop GPU (4 GB)
NVIDIA · workstation
4 GB
-60.5 GB

Doesn't fit a 24 GB card? Try this.

Practical workarounds, ranked by ease.

Drop to Q2_K
At Q2_K, weights become ~24.6 GB. Smallest quality hit per VRAM saved.
Reduce context to 4K
KV cache scales linearly with tokens. Going from 32K to 4K cuts cache ~8x.
CPU + GPU offload (llama.cpp)
Keep N layers on GPU, spill the rest to system RAM. Slower tokens/sec but it runs anywhere.
Apple Silicon unified memory
An M3 Ultra with 192 GB unified memory can hold models far larger than any single consumer GPU.
Multi-GPU tensor parallel
Two RTX 3090s = 48 GB usable with NVLink. vLLM and ExLlamaV2 support this natively.
Switch to Q8 KV cache
Halves KV memory with almost no quality loss. Supported in llama.cpp, vLLM, ExLlamaV2.

How VRAM is consumed by LLMs

The total number you see above is the sum of four components. Knowing them lets you tune for any GPU.

1. Model weights

The weights themselves. parametersB × bytesPerParam[quantization] × 1.10 framework overhead. This is the biggest term and the only one you control with quantization.

W = N × B(q) × 1.10

2. KV cache

During generation, each token's attention keys and values are cached. Grows linearly with context length and batch size. 2 × layers × hidden × bytes per token.

KV = 2 × L × H × T × batch

3. Activations

Forward-pass intermediates. Flash Attention keeps this small in modern frameworks, we budget a 1 GB floor or 5% of weights, whichever is larger.

A = max(1 GB, 0.05 × W)

4. Safety margin

CUDA kernel cache, allocator fragmentation, OS reservation, runtime buffers. Already rolled into the overhead term, never run a model at 100% of your VRAM.

leave 5–10% free

Quantization explained

Each step down halves precision roughly. Quality loss is small until you cross the Q4 cliff.

QuantBits per weightQuality lossNotes
FP16160%Reference precision. Datacenter inference.
Q8_08<0.1%Effectively lossless. The safe quantization.
Q6_K6.6~0.2%Very high quality. Great if you can afford the VRAM.
Q5_K_M5.7~0.5%Excellent balance. Recommended above Q4 when room allows.
Q4_K_M4.85~1%Sweet spot. Most popular GGUF quant. Highly recommended.
Q4_04.55~2%Legacy 4-bit. Use Q4_K_M instead when available.
Q3_K_M3.9~4%Noticeably degraded but still usable. Last resort for VRAM-starved setups.
Q2_K3~10%Significant degradation. Only when nothing else fits.
INT88<0.5%Datacenter-format 8-bit. Used by vLLM, TensorRT-LLM.
INT44~2%Datacenter-format 4-bit. AWQ, GPTQ, smoothquant.
Workaround

When to use CPU offloading

llama.cpp lets you keep a portion of model layers on the GPU and spill the rest to system RAM. The CPU computes those layers, which is dramatically slower, typically 5-15x below a fully-resident GPU run.

  • Use the -ngl N flag in llama.cpp to choose how many layers stay on the GPU.
  • Push -ngl as high as it goes without OOM. Each additional layer is a measurable speedup.
  • DDR5-6400 dual-channel pushes ~80 GB/s. That's the ceiling for CPU-resident layers, vs ~1000 GB/s on an RTX 4090.
  • Prefer offloading for batch=1 chat. For batched serving, multi-GPU beats CPU offload.

Example: Llama 3 70B Q4_K_M on RTX 4090 + 64 GB DDR5

Approximate speeds. Real numbers depend on prompt length and batch shape.

-ngl 80 (all on GPU)
OOM,
-ngl 60 (75% on GPU)
OK~9 tok/s
-ngl 40 (50% on GPU)
OK~6 tok/s
-ngl 20 (25% on GPU)
OK~3.5 tok/s
-ngl 0 (CPU only)
OK~2 tok/s

Tip: set -fa for Flash Attention and --cache-type-k q8_0 to cut KV memory in half.

Apple Silicon unified memory

When VRAM and system RAM are the same pool, the calculator's rules change.

The good news

An M3 Ultra Mac Studio with 192 GB unified memory can hold a 70B model at FP16 (~154 GB), something no single consumer GPU can do. The GPU sees the full RAM as addressable VRAM.

The catch: bandwidth

M3 Ultra: ~800 GB/s. RTX 4090: ~1000 GB/s. M3 Max: ~400 GB/s. A 4090 with 24 GB will outpace an M3 Max on anything that fits in 24 GB. Apple wins on capacity, not throughput.

How to use it

Use llama.cpp Metal backend, or MLX. The OS reserves some memory for itself, practical ceiling is ~75% of total RAM for model weights + KV cache. Plan for that headroom.

Rule of thumb: Apple's unified memory wins when you need to hold a model bigger than any consumer GPU offers (>24 GB). For models that fit in a 4090, the 4090 will be faster.

Frequently asked questions

Sanity checks, edge cases, and how the math actually works.

It's a close engineering estimate. The model weight number is exact for a given quantization; the KV cache estimate assumes standard multi-head attention and slightly overestimates models that use GQA (Llama 3, Qwen 2.5, etc.), which is the safe direction. We add a 10% framework overhead and a 1 GB activation floor. Real-world usage typically lands within ±10% of our number.

Now find the right hardware.

You know how much VRAM you need. Next step: pick the GPU, benchmark its real-world tokens-per-second, or design a full build.

Formulas based on public documentation from Hugging Face, llama.cpp, and vLLM.

Calibrated against measured llama.cpp, vLLM, and ExLlamaV2 workloads. Last updated 2025.

Spot an error? Tell us.