BGEEmbeddingMITSep 2023

BGE-Large EN v1.5

BGE-Large EN v1.5 from BAAI is the reference English embedding model for RAG pipelines — produces 1024-dim vectors, scores near the top of MTEB retrieval benchmarks. The tiny 335 M footprint runs on CPU comfortably or 100k+ embeddings/sec on a single RTX 4090. The 512-token context cap is the main limitation versus Nomic Embed v1.5.

Parameters
0.335B
dense
Context
512
tokens
Min VRAM (Q4)
0.2 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run BGE-Large EN v1.5?

BGE-Large EN v1.5 needs at minimum 0.2 GB of VRAM at Q4_K_M quantization (0.7 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A380 6GB (6 GB VRAM, $139 MSRP). Measured throughput hits 23500 emb/s on NVIDIA B200 192GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: BGE-Large EN v1.5 (BGE, 0.335B params)As of 2023-09-12

TL;DR, what to buy

Recommended GPU
Intel Arc A380 6GB
6 GB VRAM · $139 MSRP
Min VRAM at Q4_K_M
0.2 GB
+ ~20-30% headroom for KV cache
Best measured speed
23500 emb/s
on NVIDIA B200 192GB

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP160.7 GBReferenceTraining-precision weights
Q8_00.4 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K0.3 GBVery high6-bit, near Q8 quality
Q5_K_M0.3 GBHighStrong middle ground
Q4_K_M0.2 GBBalanced (recommended)Default for local deployments
Q4_00.2 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M0.2 GBLossyWhen VRAM is very tight
Q2_K0.1 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 0.7 GB; Q4_K_M needs 0.2 GB.

At full precision (FP16)

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

31 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
NVIDIA B200 192GB23500 emb/sFP165121000W
NVIDIA H200 141GB15200 emb/sFP16512700W
AMD Instinct MI300X 192GB14800 emb/sFP16512750W
NVIDIA H100 SXM5 80GB12400 emb/sFP16512700W
NVIDIA GeForce RTX 5090 32GB12200 emb/sFP16512575W
NVIDIA L40S 48GB10500 emb/sFP16512350W
AMD Instinct MI300X 192GB9800 emb/sFP16512750W
AMD Instinct MI300X 192GB9600 emb/sFP16512750W
NVIDIA A100 SXM4 80GB9500 emb/sFP16512400W
NVIDIA A100 80GB SXM9100 emb/sFP16512400W
NVIDIA GeForce RTX 5090 32GB9100 emb/sFP16512575W
NVIDIA GeForce RTX 4090 24GB8500 emb/sFP16512450W
NVIDIA GeForce RTX 5090 32GB8200 emb/sFP16512575W
NVIDIA GeForce RTX 5080 16GB8200 emb/sFP16512360W
NVIDIA RTX 6000 Ada 48GB8200 emb/sFP16512300W
NVIDIA L40S 48GB7400 emb/sFP16512350W
NVIDIA GeForce RTX 5080 16GB7200 emb/sFP16512360W
NVIDIA GeForce RTX 5070 Ti 16GB7100 emb/sFP16512300W
NVIDIA GeForce RTX 4080 Super 16GB6900 emb/sFP16512320W
NVIDIA GeForce RTX 5080 16GB6800 emb/sFP16512360W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face
Ollama
One-line install
ollama pull bge-large

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{bge-large-en-2023,
  title={BGE-Large EN v1.5},
  author={BAAI},
  year={2023},
  url={https://huggingface.co/BAAI/bge-large-en-v1.5}
}
APA
BAAI (2023). BGE-Large EN v1.5 [Model card]. Hugging Face. https://huggingface.co/BAAI/bge-large-en-v1.5
Plain text
BGE-Large EN v1.5 (BGE, BAAI, 2023). Available at https://huggingface.co/BAAI/bge-large-en-v1.5.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models