MistralText LLMApache 2.0Sep 2023

Mistral 7B

Mistral 7B v0.3 remains the reference small open model — first to popularize sliding-window attention and grouped-query attention. It's been surpassed on raw benchmarks by Llama 3.1 8B and Qwen 2.5 7B, but its Apache 2.0 license and massive ecosystem of fine-tunes keep it relevant. Excellent base for custom fine-tunes.

Parameters
7B
dense
Context
32K
tokens
Min VRAM (Q4)
4.2 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run Mistral 7B?

Mistral 7B needs at minimum 4.2 GB of VRAM at Q4_K_M quantization (15.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A380 6GB (6 GB VRAM, $139 MSRP). Measured throughput hits 1850 tok/s on Cerebras WSE-3. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: Mistral 7B (Mistral, 7B params)As of 2023-09-27

TL;DR, what to buy

Recommended GPU
Intel Arc A380 6GB
6 GB VRAM · $139 MSRP
Min VRAM at Q4_K_M
4.2 GB
+ ~20-30% headroom for KV cache
Best measured speed
1850 tok/s
on Cerebras WSE-3

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP1615.4 GBReferenceTraining-precision weights
Q8_07.7 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K6.2 GBVery high6-bit, near Q8 quality
Q5_K_M5.3 GBHighStrong middle ground
Q4_K_M4.2 GBBalanced (recommended)Default for local deployments
Q4_04.3 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M3.3 GBLossyWhen VRAM is very tight
Q2_K2.5 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 15.4 GB; Q4_K_M needs 4.2 GB.

At full precision (FP16)

  • Intel Arc Pro B5016 GB · $299
  • AMD RX 7600 XT16 GB · $329
  • Intel Arc A770 16GB16 GB · $329
  • AMD RX 9060 XT 16GB16 GB · $349
  • Intel Arc B770 16GB [VERIFY]16 GB · $349
  • NVIDIA RTX 5060 Ti 16GB16 GB · $429
  • NVIDIA RTX 4060 Ti 16GB16 GB · $499
  • AMD RX 7800 XT16 GB · $499
  • Intel Arc Pro B6024 GB · $500
  • AMD RX 907016 GB · $549

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

98 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
Cerebras WSE-31850 tok/sFP16409623000W
Cerebras WSE-3 (CS-3)1620 tok/sFP16819223000W
Groq LPU (cloud)1280 tok/sFP168192350W
Groq LPU Inference Engine750 tok/sFP84096215W
Groq LPU (per chip)540 tok/sFP168192215W
NVIDIA B200 192GB415 tok/sQ4_K_M40961000W
Google TPU v5p (Trillium)410 tok/sINT84096300W
NVIDIA H200 141GB305 tok/sQ4_K_M8192700W
NVIDIA GH200 480GB305 tok/sQ4_K_M40961000W
NVIDIA H100 SXM5 80GB295 tok/sQ4_K_M16384700W
AWS Trainium2 Trn2280 tok/sINT84096350W
AMD Instinct MI300X 192GB268 tok/sQ4_K_M4096750W
NVIDIA H200 141GB268 tok/sQ4_K_M4096700W
NVIDIA H100 SXM5 80GB215 tok/sQ4_K_M4096700W
NVIDIA GeForce RTX 5090 32GB198 tok/sQ4_K_M4096575W
AMD Instinct MI300X 192GB195 tok/sQ4_K_M4096750W
Google TPU v5e195 tok/sINT84096170W
NVIDIA GeForce RTX 5090 32GB182 tok/sQ4_K_M8192575W
NVIDIA GeForce RTX 5090 32GB178 tok/sQ4_K_M4096575W
NVIDIA A100 SXM4 80GB168 tok/sQ4_K_M4096400W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull mistral:7b
LM Studio
GGUF format
lms get mistralai/Mistral-7B-Instruct-v0.3-GGUF

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{mistral-7b-2023,
  title={Mistral 7B},
  author={Mistral AI},
  year={2023},
  url={https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3}
}
APA
Mistral AI (2023). Mistral 7B [Model card]. Hugging Face. https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3
Plain text
Mistral 7B (Mistral, Mistral AI, 2023). Available at https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models