microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Head-to-head
Dedicated comparison page for two real hardware profiles. This is built from your benchmark database, trust metadata, and buyer flow instead of generic spec-sheet comparisons.
Device profile
Compared here against NVIDIA H200 141GB.
520 tok/s
192 GB
1000 W
9.5
Device profile
405B-class models at FP8 with context headroom — the capacity play for frontier open-weight models
540 tok/s
141 GB
700 W
9.4
NVIDIA B200 192GB scores 9.5 while NVIDIA H200 141GB scores 9.4 on MyAIHardware's composite rating.
NVIDIA B200 192GB wins VRAM capacity.
NVIDIA H200 141GB is the lower-power path.
On shared workload evidence, Llama 3 8B Q4 is benchmarked at 360 tok/s for NVIDIA B200 192GB and 215 tok/s for NVIDIA H200 141GB.
| Metric | NVIDIA B200 192GB | NVIDIA H200 141GB |
|---|---|---|
| MyAI rating | 9.5 | 9.4 |
| Best LLM | 520 tok/s | 540 tok/s |
| VRAM | 192 GB | 141 GB |
| TDP | 1000 W | 700 W |
| MSRP | $40k | $30k |
| Workloads | 12 | 11 |
Llama 3 8B Q4
Llama 3 8B Q4
NVIDIA B200 192GB
360 tok/s
Q4_K_M / 192GB / 2025-03-20
NVIDIA H200 141GB
215 tok/s
Q4_K_M / 141GB / 2025-02-04
Mistral 7B Q4
Mistral 7B
NVIDIA B200 192GB
415 tok/s
Q4_K_M / 192GB / 2025-01-08
NVIDIA H200 141GB
305 tok/s
Q4_K_M / 141GB / 2026-01-28
Gemma 2 9B Q4
Gemma 2 9B
NVIDIA B200 192GB
332 tok/s
Q4_K_M / 192GB / 2025-02-12
NVIDIA H200 141GB
220 tok/s
Q4_K_M / 141GB / 2024-07-30
DeepSeek-R1 Distill 7B
DeepSeek-R1 7B
NVIDIA B200 192GB
480 tok/s
Q4_K_M / 192GB / 2026-02-04
NVIDIA H200 141GB
305 tok/s
Q4_K_M / 141GB / 2026-04-20
Qwen 2.5 14B Q4
Qwen 2.5 14B
NVIDIA B200 192GB
285 tok/s
Q4_K_M / 192GB / 2026-01-15
NVIDIA H200 141GB
178 tok/s
Q4_K_M / 141GB / 2024-12-15
Llama 3 70B Q4
Llama 3 70B Q4
NVIDIA B200 192GB
138 tok/s
Q4_K_M / 192GB / 2024-12-26
NVIDIA H200 141GB
84.0 tok/s
Q4_K_M / 141GB / 2024-04-06
This shortlist cross-references the external OpenEvals snapshot with an approximate Q4-class VRAM estimate. Use it to sanity-check whether the hardware you are comparing can host strong open models, not just benchmark toy workloads.
microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Qwen/Qwen2-72B
73B params · est. 43.6 GB Q4
microsoft/Phi-3.5-mini-instruct
3.8B params · est. 2.3 GB Q4
internlm/internlm2_5-7b-chat
7.7B params · est. 4.6 GB Q4