microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Head-to-head
Dedicated comparison page for two real hardware profiles. This is built from your benchmark database, trust metadata, and buyer flow instead of generic spec-sheet comparisons.
Device profile
405B-class models at FP8 with context headroom — the capacity play for frontier open-weight models
540 tok/s
141 GB
700 W
9.4
Device profile
70B FP16 with substantial context, 405B Q4 with headroom. Production multi-tenant serving at batch 8-64 via TensorRT-LLM or vLLM.
612 tok/s
80 GB
700 W
9.4
NVIDIA H200 141GB scores 9.4 while NVIDIA H100 SXM5 80GB scores 9.4 on MyAIHardware's composite rating.
NVIDIA H200 141GB wins VRAM capacity.
Power draw is effectively tied on paper.
On shared workload evidence, Llama 3 8B Q4 is benchmarked at 215 tok/s for NVIDIA H200 141GB and 178 tok/s for NVIDIA H100 SXM5 80GB.
| Metric | NVIDIA H200 141GB | NVIDIA H100 SXM5 80GB |
|---|---|---|
| MyAI rating | 9.4 | 9.4 |
| Best LLM | 540 tok/s | 612 tok/s |
| VRAM | 141 GB | 80 GB |
| TDP | 700 W | 700 W |
| MSRP | $30k | $25k |
| Workloads | 11 | 13 |
Llama 3 8B Q4
Llama 3 8B Q4
NVIDIA H200 141GB
215 tok/s
Q4_K_M / 141GB / 2025-02-04
NVIDIA H100 SXM5 80GB
178 tok/s
Q4_K_M / 80GB / 2024-11-08
Mistral 7B Q4
Mistral 7B
NVIDIA H200 141GB
305 tok/s
Q4_K_M / 141GB / 2026-01-28
NVIDIA H100 SXM5 80GB
295 tok/s
Q4_K_M / 80GB / 2025-05-08
Gemma 2 9B Q4
Gemma 2 9B
NVIDIA H200 141GB
220 tok/s
Q4_K_M / 141GB / 2024-07-30
NVIDIA H100 SXM5 80GB
215 tok/s
Q4_K_M / 80GB / 2024-10-14
DeepSeek-R1 Distill 7B
DeepSeek-R1 7B
NVIDIA H200 141GB
305 tok/s
Q4_K_M / 141GB / 2026-04-20
NVIDIA H100 SXM5 80GB
245 tok/s
Q4_K_M / 80GB / 2024-08-04
Qwen 2.5 14B Q4
Qwen 2.5 14B
NVIDIA H200 141GB
178 tok/s
Q4_K_M / 141GB / 2024-12-15
NVIDIA H100 SXM5 80GB
168 tok/s
Q4_K_M / 80GB / 2024-10-21
Llama 3 70B Q4
Llama 3 70B Q4
NVIDIA H200 141GB
84.0 tok/s
Q4_K_M / 141GB / 2024-04-06
NVIDIA H100 SXM5 80GB
66.0 tok/s
Q4_K_M / 80GB / 2024-06-24
This shortlist cross-references the external OpenEvals snapshot with an approximate Q4-class VRAM estimate. Use it to sanity-check whether the hardware you are comparing can host strong open models, not just benchmark toy workloads.
microsoft/Phi-3-medium-4k-instruct
14B params · est. 8.4 GB Q4
Qwen/Qwen2-72B
73B params · est. 43.6 GB Q4
microsoft/Phi-3.5-mini-instruct
3.8B params · est. 2.3 GB Q4
internlm/internlm2_5-7b-chat
7.7B params · est. 4.6 GB Q4