Datacenter GPUNVIDIA

NVIDIA A100 80GB SXM

Curated Aggregate·2024-08-287 workloads · 8 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

80 GB

TDP

400 W

MSRP

$15k

Perf/W

0.12 tok/s/W

Cost/1K tok

$0.0033/k

Tested

2024-08-28

Quick answer

How many tokens per second does NVIDIA A100 80GB SXM produce on Llama 3 70B Q4?

NVIDIA A100 80GB SXM has a source-attributed result of 48.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-a100-80gb-l3-70b-q4)As of 2024-08-28

Overview

Cloud Standard

NVIDIA A100 SXM4 80GB — Ampere-architecture datacenter GPU from 2020. 80GB HBM2e at 2 TB/s, 400W TDP. The workhorse that trained the GPT-3 generation of models. Still widely deployed in cloud and on-prem.

AI Usefulness

80GB supports many 70B quantized configurations. 70B FP16 weights need about 140GB before cache and runtime overhead, so they require more than one device or explicit offload. Use matched model settings to compare published throughput.

Verdict

NVIDIA A100 80GB SXM with 80GB VRAM at 400W TDP, scored across 7 workloads with 8 benchmark records.

Reference workload

Asure 12B

145 tok/s

Quantization

Q4_K_M

8K context · batch 1

LLM Inference Performance

055110165220Llama 370B Q4Llama 3 8BFP16Mistral 7BGemma 29BQwen 2.514BAsure 12B

Benchmarks (7 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

48.0tok/sQ4_K_M4K, Curated Aggregate2024-08-28
Llama 3 8B FP16

llm

215.0tok/sFP164K, Curated Aggregate2024-09-14
Embedding throughput

embedding

9100.0emb/sFP161K, Curated Aggregate2024-09-25
Mistral 7B

llm

168.0tok/sQ4_K_M4K, Curated Aggregate2024-04-15
Gemma 2 9B

llm

138.0tok/sQ4_K_M8K, Curated Aggregate2024-07-09
Qwen 2.5 14B

llm

108.0tok/sQ4_K_M8K, Curated Aggregate2024-10-30
Asure 12B

llm

145.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA A100 80GB SXM. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 80GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Good
70B
Excellent
Tight
Won't fit
Reported batch-one examples
Asure 12B145 tok/s
Qwen 2.5 14B108 tok/s
Embedding throughput9100 emb/s

Source

Trendyol-LLM-Asure 12B on A100.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2025-05-01; 496 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $15k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.