BGE-Large Embeddings
BAAI/bge-large-en-v1.5, FP16, 512-token inputs, batch 32.
Primary metric: Embeddings / sec (emb/s)
Reference prompts
Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.
- Prompt 1
512-token Wikipedia paragraph (English).
- Prompt 2
512-token code comment block (Python).
- Prompt 3
512-token product description (e-commerce).
Reference runtime command
A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.
text-embeddings-router --model-id BAAI/bge-large-en-v1.5 --dtype float16 --max-batch-tokens 16384Batch 32 of 512-token inputs. Sustained throughput, excluding cold-start. We report embeddings/sec.
Full leaderboard
Every record for Embedding throughput.
Top 10 leaderboard
BGE-Large Embeddings · sorted by embeddings / sec
Value frontier, MSRP vs embeddings / sec
Each dot is a device. Top-left is best value (cheap + fast).
Full leaderboard
Click any row for detailed breakdown. Click column headers to sort.
| # | Device | Verif | Buy | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | NVIDIA B200 192GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 23.5kemb/s | 192 GB | 1000 W | $40k | — | — | Amazon | |
| 2 | NVIDIA H200 141GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 15.2kemb/s | 141 GB | 700 W | $30k | — | — | Amazon | |
| 3 | AMD Instinct MI300X 192GBDatacenter GPU AMD·FP16·1K ctx | 6 | 14.8kemb/s | 192 GB | 750 W | $15k | — | — | Amazon | |
| #4 | NVIDIA H100 SXM5 80GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 12.4kemb/s | 80 GB | 700 W | $25k | — | — | Amazon | |
| #5 | NVIDIA GeForce RTX 5090 32GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 12.2kemb/s | 32 GB | 575 W | $2.0k | — | — | Amazon | |
| #6 | NVIDIA L40S 48GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 10.5kemb/s | 48 GB | 350 W | $7.8k | — | — | Amazon | |
| #7 | AMD Instinct MI300X 192GBDatacenter GPU AMD·FP16·1K ctx | 6 | 9.8kemb/s | 192 GB | 750 W | $15k | — | — | Amazon | |
| #8 | AMD Instinct MI300X 192GBDatacenter GPU AMD·FP16·1K ctx | 6 | 9.6kemb/s | 192 GB | 750 W | $18k | — | — | Amazon | |
| #9 | NVIDIA A100 SXM4 80GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 9.5kemb/s | 80 GB | 400 W | $15k | — | — | Amazon | |
| #10 | NVIDIA A100 80GB SXMDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 9.1kemb/s | 80 GB | 400 W | $15k | — | — | Amazon | |
| #11 | NVIDIA GeForce RTX 5090 32GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 9.1kemb/s | 32 GB | 575 W | $2.0k | — | — | Amazon | |
| #12 | NVIDIA GeForce RTX 4090 24GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 8.5kemb/s | 24 GB | 450 W | $1.6k | — | — | Amazon | |
| #13 | NVIDIA GeForce RTX 5090 32GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 8.2kemb/s | 32 GB | 575 W | $2.0k | — | — | Amazon | |
| #14 | NVIDIA GeForce RTX 5080 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 8.2kemb/s | 16 GB | 360 W | $999 | — | — | Amazon | |
| #15 | NVIDIA RTX 6000 Ada 48GBPro GPU NVIDIA·FP16·1K ctx | 6 | 8.2kemb/s | 48 GB | 300 W | $6.8k | — | — | Amazon | |
| #16 | NVIDIA L40S 48GBDatacenter GPU NVIDIA·FP16·1K ctx | 6 | 7.4kemb/s | 48 GB | 350 W | $7.8k | — | — | Amazon | |
| #17 | NVIDIA GeForce RTX 5080 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 7.2kemb/s | 16 GB | 360 W | $999 | — | — | Amazon | |
| #18 | NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 7.1kemb/s | 16 GB | 300 W | $749 | — | — | Amazon | |
| #19 | NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 6.9kemb/s | 16 GB | 320 W | $999 | — | — | Amazon | |
| #20 | NVIDIA GeForce RTX 5080 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 6.8kemb/s | 16 GB | 360 W | $999 | — | — | Amazon | |
| #21 | NVIDIA GeForce RTX 4090 24GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 5.8kemb/s | 24 GB | 450 W | $1.6k | — | — | Amazon | |
| #22 | AMD Radeon RX 7900 XTX 24GBConsumer GPU AMD·FP16·1K ctx | 6 | 5.8kemb/s | 24 GB | 355 W | $999 | — | — | Amazon | |
| #23 | NVIDIA GeForce RTX 4090 24GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 5.4kemb/s | 24 GB | 450 W | $1.6k | — | — | Amazon | |
| #24 | NVIDIA GeForce RTX 3090 24GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 4.9kemb/s | 24 GB | 350 W | $1.5k | — | — | Amazon | |
| #25 | Apple M3 Ultra (80c GPU, 512GB)Apple Silicon Apple·FP16·1K ctx | 6 | 4.9kemb/s | 512 GB | 100 W | $10.0k | — | — | Amazon | |
| #26 | NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 4.2kemb/s | 16 GB | 320 W | $999 | — | — | Amazon | |
| #27 | Apple M4 Max (40c GPU, 128GB)Apple Silicon Apple·FP16·1K ctx | 6 | 3.5kemb/s | 128 GB | 65 W | $4.7k | — | — | Amazon | |
| #28 | NVIDIA GeForce RTX 3090 24GBConsumer GPU NVIDIA·FP16·1K ctx | 6 | 3.2kemb/s | 24 GB | 350 W | $1.5k | — | — | Amazon | |
| #29 | Apple M3 Ultra (80c GPU, 512GB)Apple Silicon Apple·FP16·1K ctx | 6 | 2.7kemb/s | 512 GB | 80 W | $8.5k | — | — | Amazon | |
| #30 | Apple M4 Max (40c GPU, 128GB)Apple Silicon Apple·FP16·1K ctx | 6 | 1.9kemb/s | 128 GB | 65 W | $5.0k | — | — | Amazon | |
| #31 | Apple M4 Max (40c GPU, 128GB)Apple Silicon Apple·FP16·1K ctx | 6 | 1.8kemb/s | 128 GB | 70 W | $4.7k | — | — | Amazon |
Cite this benchmark
Use this in your paper, blog post, or comparison table.
@misc{myaihardware_embedding-bge-large_2026,
title = {MyAI Bench: BGE-Large Embeddings},
author = {{MyAIHardware Contributors}},
year = {2026},
url = {https://www.myaihardware.com/benchmarks/workload/embedding-bge-large},
note = {Version 1.3, accessed 2026-08-27}
}MyAIHardware Contributors. (2026). MyAI Bench: BGE-Large Embeddings. MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/embedding-bge-large
MyAIHardware Contributors. "MyAI Bench: BGE-Large Embeddings." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/embedding-bge-large. Accessed 2026-08-27.
MyAI Bench, BGE-Large Embeddings. MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/embedding-bge-large (accessed 2026-08-27).