Datacenter GPUNVIDIA

NVIDIA A100 80GB SXM

Curated Aggregate·2024-08-287 workloads · 8 records
MyAI Rating9.0emb/s · Embedding throughput

VRAM

80 GB

TDP

400 W

MSRP

$15k

Perf/W

0.12 emb/s/W

Cost/1K tok

$0.0033/k

Tested

2024-08-28

Quick answer

How many tokens per second does NVIDIA A100 80GB SXM produce on Llama 3 70B Q4?

NVIDIA A100 80GB SXM produces approximately 48.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 80GB VRAM, 400W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-a100-80gb-l3-70b-q4)As of 2024-08-28

Overview

Cloud Standard

NVIDIA A100 SXM4 80GB — Ampere-architecture datacenter GPU from 2020. 80GB HBM2e at 2 TB/s, 400W TDP. The workhorse that trained the GPT-3 generation of models. Still widely deployed in cloud and on-prem.

AI Usefulness

80GB VRAM hosts 70B FP16 with context headroom. The most widely available datacenter GPU for inference — every cloud provider has A100 instances. ~145 tok/s on 12B Q4. Mature software support across every engine and framework. Best for: cloud-based inference, training runs up to 70B, and organizations with existing A100 infrastructure. Being displaced by H100/H200 for new deployments but remains the most accessible datacenter tier.

Verdict

NVIDIA A100 80GB SXM with 80GB VRAM at 400W TDP, scored across 7 workloads with 8 benchmark records.

Best workload

Embedding throughput

9500 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

055110165220Llama 370B Q4Llama 38B FP16Mistral 7BGemma 29BQwen 2.514BAsure 12B

Benchmarks (7 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

48.0tok/sQ4_K_M4K, Curated Aggregate2024-08-28
Llama 3 8B FP16

llm

215.0tok/sFP164K, Curated Aggregate2024-09-14
Embedding throughput

embedding

9500.0emb/sFP161K, Curated Aggregate2024-05-21
Mistral 7B

llm

168.0tok/sQ4_K_M4K, Curated Aggregate2024-04-15
Gemma 2 9B

llm

138.0tok/sQ4_K_M8K, Curated Aggregate2024-07-09
Qwen 2.5 14B

llm

108.0tok/sQ4_K_M8K, Curated Aggregate2024-10-30
Asure 12B

llm

145.0tok/sQ4_K_M8K±5.8Curated Aggregate2025-05-01

MyAI Score

Reference
9.0/10

NVIDIA A100 80GB SXM clears a 9.0/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
307
Capability
185
Efficiency
52
Value
27
Trust
44
Coverage
45
Composite benchmark660 / 1000

Workload Fit

What models fit this 80GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Good
70B
Excellent
Good
Won't fit
Top Benchmarks
Embedding throughput9500 emb/s
Llama 3 8B FP16215 tok/s
Mistral 7B168 tok/s

Public Trust Layer

Trust score

6/10

MyAI rating

9.0

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-05-21; 828 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $15k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.