Datacenter GPUNVIDIA

NVIDIA B200 192GB

Vendor Claim·2024-12-2612 workloads · 14 records
MyAI Rating9.5emb/s · Embedding throughput

VRAM

192 GB

TDP

1000 W

MSRP

$40k

Perf/W

0.14 emb/s/W

Cost/1K tok

$0.0031/k

Tested

2024-12-26

Quick answer

How many tokens per second does NVIDIA B200 192GB produce on Llama 3 70B Q4?

NVIDIA B200 192GB produces approximately 138.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 8192-token context, 192GB VRAM, 1000W TDP). That figure comes from 4 measured runs on TensorRT-LLM 0.15. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-b200-l3-70b-q4)As of 2024-12-26

Overview

Frontier Datacenter

NVIDIA B200 192GB — Blackwell-architecture datacenter GPU. 192GB HBM3e at 8 TB/s, 1000W TDP. Second-gen Transformer Engine with FP4 support. 2.5x the AI compute of H100. NVIDIA's 2025 flagship datacenter GPU.

AI Usefulness

192GB VRAM at 8 TB/s bandwidth makes this the fastest single-GPU inference platform available. Hosts 405B at FP8 with massive context headroom. FP4 doubles effective throughput when engines support it. Currently the ceiling for single-GPU local AI — everything above requires tensor parallelism. Priced for cloud/enterprise, not homelab.

Verdict

NVIDIA B200 192GB with 192GB VRAM at 1000W TDP, scored across 12 workloads with 14 benchmark records.

Best workload

Embedding throughput

23500 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

0150300450600Llama 370B Q4Llama 370B Q8Llama 38B FP16Qwen 2.514BDeepSeek-R17BMistral 7B

Benchmarks (12 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

138.0tok/sQ4_K_M8K±12.4Vendor Claim2024-12-26
Llama 3 70B Q8

llm

98.0tok/sQ8_08K, Curated Aggregate2024-12-30
Llama 3 8B FP16

llm

520.0tok/sFP168K, Curated Aggregate2026-02-04
SDXL image gen

image

52.0img/minFP16, , Curated Aggregate2025-04-15
Qwen 2.5 14B

llm

285.0tok/sQ4_K_M8K, Curated Aggregate2026-01-15
DeepSeek-R1 7B

llm

480.0tok/sQ4_K_M8K, Curated Aggregate2026-02-04
Whisper transcription

audio

210.0x RTFP16, , Curated Aggregate2026-02-22
Mistral 7B

llm

415.0tok/sQ4_K_M4K, Curated Aggregate2025-01-08
Gemma 2 9B

llm

332.0tok/sQ4_K_M8K, Curated Aggregate2025-02-12
Llama 3 8B Q4

llm

360.0tok/sQ4_K_M128K, Curated Aggregate2025-03-20
Embedding throughput

embedding

23500.0emb/sFP161K, Curated Aggregate2025-01-18
Asure 12B

llm

350.0tok/sQ4_K_M8K±10.5Vendor Claim2025-05-01

MyAI Score

Mythic
9.5/10

NVIDIA B200 192GB clears a 9.5/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
384
Capability
197
Efficiency
57
Value
29
Trust
48
Coverage
60
Composite benchmark775 / 1000

Workload Fit

What models fit this 192GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Excellent
Top Benchmarks
Embedding throughput23500 emb/s
Llama 3 8B FP16520 tok/s
DeepSeek-R1 7B480 tok/s

Source

TEI 1.5 on Blackwell DC.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

9.5

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2025-01-18; 586 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $40k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.