embedding workload · 31 runs on record

BGE-Large Embeddings

BAAI/bge-large-en-v1.5, FP16, 512-token inputs, batch 32.

Primary metric: Embeddings / sec (emb/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    512-token Wikipedia paragraph (English).

  • Prompt 2

    512-token code comment block (Python).

  • Prompt 3

    512-token product description (e-commerce).

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
text-embeddings-router --model-id BAAI/bge-large-en-v1.5 --dtype float16 --max-batch-tokens 16384

Batch 32 of 512-token inputs. Sustained throughput, excluding cold-start. We report embeddings/sec.

Full leaderboard

Every record for Embedding throughput.

Top 10 leaderboard

BGE-Large Embeddings · sorted by embeddings / sec

Bar chart: Top 10 devices ranked by Embeddings / sec for BGE-Large Embeddings. 1. NVIDIA B200 192GB at 23,500 emb/s. 2. NVIDIA H200 141GB at 15,200 emb/s. 3. AMD Instinct MI300X 192GB at 14,800 emb/s. 4. NVIDIA H100 SXM5 80GB at 12,400 emb/s. 5. NVIDIA GeForce RTX 5090 32GB at 12,200 emb/s.

Value frontier, MSRP vs embeddings / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Embeddings / sec across 31 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: B200 192GB at $39,999 delivering 23,500 emb/s; H200 141GB at $30,000 delivering 15,200 emb/s; Instinct MI300X 192GB at $15,000 delivering 14,800 emb/s; H100 SXM5 80GB at $25,000 delivering 12,400 emb/s; GeForce RTX 5090 32GB at $1,999 delivering 12,200 emb/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
NVIDIA B200 192GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
23.5kemb/s
192 GB1000 W$40kAmazon
2
NVIDIA H200 141GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
15.2kemb/s
141 GB700 W$30kAmazon
3
AMD Instinct MI300X 192GBDatacenter GPU
AMD·FP16·1K ctx
6
14.8kemb/s
192 GB750 W$15kAmazon
#4
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
12.4kemb/s
80 GB700 W$25kAmazon
#5
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·FP16·1K ctx
6
12.2kemb/s
32 GB575 W$2.0kAmazon
#6
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
10.5kemb/s
48 GB350 W$7.8kAmazon
#7
AMD Instinct MI300X 192GBDatacenter GPU
AMD·FP16·1K ctx
6
9.8kemb/s
192 GB750 W$15kAmazon
#8
AMD Instinct MI300X 192GBDatacenter GPU
AMD·FP16·1K ctx
6
9.6kemb/s
192 GB750 W$18kAmazon
#9
NVIDIA A100 SXM4 80GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
9.5kemb/s
80 GB400 W$15kAmazon
#10
NVIDIA A100 80GB SXMDatacenter GPU
NVIDIA·FP16·1K ctx
6
9.1kemb/s
80 GB400 W$15kAmazon
#11
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·FP16·1K ctx
6
9.1kemb/s
32 GB575 W$2.0kAmazon
#12
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·FP16·1K ctx
6
8.5kemb/s
24 GB450 W$1.6kAmazon
#13
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·FP16·1K ctx
6
8.2kemb/s
32 GB575 W$2.0kAmazon
#14
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
8.2kemb/s
16 GB360 W$999Amazon
#15
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·FP16·1K ctx
6
8.2kemb/s
48 GB300 W$6.8kAmazon
#16
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·FP16·1K ctx
6
7.4kemb/s
48 GB350 W$7.8kAmazon
#17
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
7.2kemb/s
16 GB360 W$999Amazon
#18
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
7.1kemb/s
16 GB300 W$749Amazon
#19
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
6.9kemb/s
16 GB320 W$999Amazon
#20
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
6.8kemb/s
16 GB360 W$999Amazon
#21
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·FP16·1K ctx
6
5.8kemb/s
24 GB450 W$1.6kAmazon
#22
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·FP16·1K ctx
6
5.8kemb/s
24 GB355 W$999Amazon
#23
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·FP16·1K ctx
6
5.4kemb/s
24 GB450 W$1.6kAmazon
#24
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·FP16·1K ctx
6
4.9kemb/s
24 GB350 W$1.5kAmazon
#25
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·FP16·1K ctx
6
4.9kemb/s
512 GB100 W$10.0kAmazon
#26
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·FP16·1K ctx
6
4.2kemb/s
16 GB320 W$999Amazon
#27
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·FP16·1K ctx
6
3.5kemb/s
128 GB65 W$4.7kAmazon
#28
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·FP16·1K ctx
6
3.2kemb/s
24 GB350 W$1.5kAmazon
#29
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·FP16·1K ctx
6
2.7kemb/s
512 GB80 W$8.5kAmazon
#30
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·FP16·1K ctx
6
1.9kemb/s
128 GB65 W$5.0kAmazon
#31
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·FP16·1K ctx
6
1.8kemb/s
128 GB70 W$4.7kAmazon
Sorted by Embeddings / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_embedding-bge-large_2026,
  title  = {MyAI Bench: BGE-Large Embeddings},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/embedding-bge-large},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: BGE-Large Embeddings. MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/embedding-bge-large
MLA
MyAIHardware Contributors. "MyAI Bench: BGE-Large Embeddings." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/embedding-bge-large. Accessed 2026-08-27.
Plain text
MyAI Bench, BGE-Large Embeddings. MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/embedding-bge-large (accessed 2026-08-27).