Pro GPUNVIDIA

NVIDIA RTX 6000 Ada 48GB

Curated Aggregate·2024-12-148 workloads · 8 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

48 GB

TDP

300 W

MSRP

$6.8k

Perf/W

0.10 tok/s/W

Cost/1K tok

$0.0024/k

Tested

2024-12-14

Quick answer

How many tokens per second does NVIDIA RTX 6000 Ada 48GB produce on Llama 3 70B Q4?

NVIDIA RTX 6000 Ada 48GB has a source-attributed result of 30.0 tok/s on Llama 3 70B Q4 (batch 1, 8192-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-rtx6000-ada-l3-70b-q4)As of 2024-12-14

Overview

Pro Workstation

NVIDIA RTX 6000 Ada 48GB — professional visualization and AI GPU. 48GB GDDR6 ECC at 960 GB/s, 300W TDP. Ada Lovelace architecture with full FP8 support. Active cooling, fits in standard workstations.

AI Usefulness

48GB can hold many 70B Q4 artifacts, with context determined by cache and runtime overhead. 32B FP16 weights need about 64GB and 70B Q8 about 70GB; neither fits fully. Confirm the exact workstation SKU and cooling requirements.

Verdict

NVIDIA RTX 6000 Ada 48GB with 48GB VRAM at 300W TDP, scored across 8 workloads with 8 benchmark records.

Reference workload

Asure 12B

95 tok/s

Quantization

Q4_K_M

8K context · batch 1

LLM Inference Performance

03570105140Llama 370B Q4Llama 3 8BFP16Mistral 7BGemma 29BQwen 2.514BAsure 12B

Benchmarks (8 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

30.0tok/sQ4_K_M8K, Curated Aggregate2024-12-14
Llama 3 8B FP16

llm

88.0tok/sFP164K, Curated Aggregate2024-11-22
SDXL image gen

image

22.0img/minFP16, , Curated Aggregate2025-01-08
Mistral 7B

llm

125.0tok/sQ4_K_M4K, Curated Aggregate2024-05-29
Gemma 2 9B

llm

104.0tok/sQ4_K_M8K, Curated Aggregate2024-08-11
Qwen 2.5 14B

llm

76.0tok/sQ4_K_M8K, Curated Aggregate2024-10-22
Embedding throughput

embedding

8200.0emb/sFP161K, Curated Aggregate2024-09-14
Asure 12B

llm

95.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA RTX 6000 Ada 48GB. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 48GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Won't fit
70B
Good
Won't fit
Won't fit
Reported batch-one examples
Asure 12B95 tok/s
SDXL image gen22 img/min
Llama 3 70B Q430 tok/s

Source

Trendyol-LLM-Asure 12B RTX 6000 Ada.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2025-05-01; 496 days old.

Open primary source

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

Amazon is usually the quickest place to sanity-check live pricing on workstation-class gear before you compare specialty retailers.

Search Amazon listings
  • +Cross-check the current ask against $6.8k and any reputable open-box or used options.
  • +Verify seller reputation, warranty terms, and the exact board configuration before you buy.
  • +If you plan long 24/7 inference runs, warranty and thermals matter more than a tiny discount.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.