PhiText LLMMITAug 2024

Phi-3.5 Mini 3.8B

Phi-3.5 Mini is Microsoft's curated-synthetic-data 3.8 B model — rivals 7-13 B models on reasoning despite the small size. Particularly notable for the 128K context window in such a small footprint. Runs at 100+ tok/s on midrange consumer GPUs at Q4_K_M (only ~2.3 GB VRAM). Weakness: limited world knowledge versus naturally-trained peers.

Parameters
3.8B
dense
Context
128K
tokens
Min VRAM (Q4)
2.3 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Phi-3.5 Mini 3.8B?

Phi-3.5 Mini 3.8B needs at minimum 2.3 GB of VRAM at Q4_K_M quantization (8.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A380 6GB (6 GB VRAM, $139 MSRP). Measured throughput hits 612 tok/s on NVIDIA H100 SXM5 80GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Phi-3.5 Mini 3.8B (Phi, 3.8B params)As of 2024-08-20

TL;DR, what to buy

Recommended GPU
Intel Arc A380 6GB
6 GB VRAM · $139 MSRP
Min VRAM at Q4_K_M
2.3 GB
+ ~20-30% headroom for KV cache
Best measured speed
612 tok/s
on NVIDIA H100 SXM5 80GB

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP168.4 GBReferenceTraining-precision weights
Q8_04.2 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K3.4 GBVery high6-bit, near Q8 quality
Q5_K_M2.9 GBHighStrong middle ground
Q4_K_M2.3 GBBalanced (recommended)Default for local deployments
Q4_02.4 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M1.8 GBLossyWhen VRAM is very tight
Q2_K1.3 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 8.4 GB; Q4_K_M needs 2.3 GB.

At full precision (FP16)

  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 2060 (12 GB refresh)12 GB · $299
  • Intel Arc Pro B5016 GB · $299
  • AMD RX 7600 XT16 GB · $329
  • Intel Arc A770 16GB16 GB · $329
  • NVIDIA GeForce RTX 3060 (12 GB)12 GB · $329
  • AMD RX 9060 XT 16GB16 GB · $349
  • Intel Arc B770 16GB [VERIFY]16 GB · $349

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

68 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
NVIDIA H100 SXM5 80GB612 tok/sQ4_K_M4096700W
NVIDIA GeForce RTX 5090 32GB348 tok/sQ4_K_M4096575W
NVIDIA GeForce RTX 5090 32GB320 tok/sQ4_K_M4096575W
Google TPU v5e320 tok/sINT84096170W
NVIDIA GeForce RTX 4090 24GB285 tok/sQ4_K_M4096450W
NVIDIA GeForce RTX 4090 24GB245 tok/sQ4_K_M4096450W
NVIDIA GeForce RTX 4080 Super 16GB215 tok/sQ4_K_M4096320W
AMD Radeon RX 7900 XTX 24GB195 tok/sQ4_K_M4096355W
AMD Instinct MI210 64GB175 tok/sQ4_K_M4096300W
NVIDIA GeForce RTX 4070 Ti 12GB165 tok/sQ4_K_M4096285W
NVIDIA GeForce RTX 5070 12GB165 tok/sQ4_K_M4096250W
NVIDIA A40 48GB162 tok/sQ4_K_M4096300W
NVIDIA GeForce RTX 5060 Ti 8GB158 tok/sQ4_K_M4096180W
NVIDIA GeForce RTX 5060 8GB145 tok/sQ4_K_M4096145W
NVIDIA GeForce RTX 4070 12GB138 tok/sQ4_K_M4096200W
NVIDIA GeForce RTX 5060 Ti 16GB132 tok/sQ4_K_M4096180W
AMD Radeon RX 7800 XT 16GB132 tok/sQ4_K_M4096263W
NVIDIA RTX A5000 24GB122 tok/sQ4_K_M4096230W
NVIDIA GeForce RTX 5060 Ti 8GB118 tok/sQ4_K_M4096165W
NVIDIA GeForce RTX 5060 8GB98 tok/sQ4_K_M4096145W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull phi3.5
LM Studio
GGUF format
lms get bartowski/Phi-3.5-mini-instruct-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{phi-3-5-mini-2024,
  title={Phi-3.5 Mini 3.8B},
  author={Microsoft Research},
  year={2024},
  url={https://huggingface.co/microsoft/Phi-3.5-mini-instruct}
}
APA
Microsoft Research (2024). Phi-3.5 Mini 3.8B [Model card]. Hugging Face. https://huggingface.co/microsoft/Phi-3.5-mini-instruct
Plain text
Phi-3.5 Mini 3.8B (Phi, Microsoft Research, 2024). Available at https://huggingface.co/microsoft/Phi-3.5-mini-instruct.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models