Datacenter GPUNVIDIA

NVIDIA L40S 48GB

Curated Aggregate·2024-06-119 workloads · 10 records
MyAI Rating9.0emb/s · Embedding throughput

VRAM

48 GB

TDP

350 W

MSRP

$7.8k

Perf/W

0.50 emb/s/W

Cost/1K tok

$0.47/M

Tested

2024-06-11

Quick answer

How many tokens per second does NVIDIA L40S 48GB produce on Llama 3 70B Q4?

NVIDIA L40S 48GB produces approximately 22.0 tok/s on Llama 3 70B at Q4_K_M quantization (batch 1, 4096-token context, 48GB VRAM, 350W TDP). That figure comes from 1 measured run on llama.cpp. The 70B model needs roughly 40GB of VRAM at Q4, so headroom and KV cache budget matter as much as raw throughput.

Source: MyAIHardware benchmark database (bench-l40s-l3-70b-q4)As of 2024-08-07

Overview

Enterprise Inference

NVIDIA L40S 48GB — Ada Lovelace datacenter GPU with 48GB GDDR6 at 864 GB/s, 350W TDP. Passive cooling (requires server airflow). Positioned between consumer cards and H100 for inference-heavy workloads.

AI Usefulness

48GB VRAM comfortably hosts 70B Q4 with large context or 32B FP16 full-GPU. The sweet spot for single-GPU enterprise inference where H100 is overkill. ~110 tok/s on 12B Q4. Passive cooling means datacenter deployment only — not a desktop card. Best for: enterprise inference serving, RAG pipelines, and multi-tenant LLM hosting where 48GB is the right size.

Verdict

NVIDIA L40S 48GB with 48GB VRAM at 350W TDP, scored across 9 workloads with 10 benchmark records.

Best workload

Embedding throughput

10500 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

04590135180Llama 38B FP16Llama 370B Q4Mistral 7BGemma 29BQwen 2.514BDeepSeek-R17B

Benchmarks (9 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B FP16

llm

175.0tok/sFP164K, Curated Aggregate2024-06-11
Llama 3 70B Q4

llm

22.0tok/sQ4_K_M4K, Curated Aggregate2024-08-07
SDXL image gen

image

18.0img/minFP16, , Curated Aggregate2026-04-29
Embedding throughput

embedding

10500.0emb/sFP161K, Curated Aggregate2024-08-04
Mistral 7B

llm

142.0tok/sQ4_K_M4K, Curated Aggregate2024-06-18
Gemma 2 9B

llm

118.0tok/sQ4_K_M8K, Curated Aggregate2024-08-23
Qwen 2.5 14B

llm

92.0tok/sQ4_K_M8K, Curated Aggregate2024-11-15
DeepSeek-R1 7B

llm

138.0tok/sQ4_K_M8K, Curated Aggregate2025-02-20
Asure 12B

llm

110.0tok/sQ4_K_M8K±5.5Curated Aggregate2025-05-01

MyAI Score

Reference
9.0/10

NVIDIA L40S 48GB clears a 9.0/10 based on workload-normalized throughput, memory headroom, efficiency, value, trust, and coverage.

Throughput
297
Capability
171
Efficiency
51
Value
32
Trust
46
Coverage
55
Composite benchmark652 / 1000

Workload Fit

What models fit this 48GB card at different quantization levels.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Won't fit
70B
Good
Won't fit
Won't fit
Top Benchmarks
Embedding throughput10500 emb/s
Llama 3 8B FP16175 tok/s
Mistral 7B142 tok/s

Public Trust Layer

Trust score

6/10

MyAI rating

9.0

Runs

1

Freshness

Stale

Source-linked row with explicit verification status.

Tested on 2024-08-04; 753 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $7.8k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

This is one of the highest-scoring parts in our database, so it is worth tracking when inventory lands near fair-market pricing.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.