ASICCerebras

Cerebras WSE-3 (CS-3)

Curated Aggregate·2025-03-294 workloads · 5 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

44 GB

TDP

23000 W

MSRP

$2.50M

Perf/W

0.02 tok/s/W

Cost/1K tok

$0.0587/k

Tested

2025-03-29

Quick answer

How many tokens per second does Cerebras WSE-3 (CS-3) produce on Llama 3 70B Q4?

Cerebras WSE-3 (CS-3) has a source-attributed result of 450.0 tok/s on Llama 3 70B Q4 (batch 1, 8192-token context, FP16; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-cerebras-l3-70b-q4)As of 2025-03-29

Overview

Wafer-Scale Frontier

Cerebras WSE-3 — wafer-scale AI accelerator with 4 trillion transistors, 900,000 AI cores, 44GB on-chip SRAM at 21 PB/s bandwidth, 23kW system power. The largest single chip ever built for AI.

AI Usefulness

Memory bandwidth (21 PB/s) is orders of magnitude beyond any GPU — token generation at 1,420 tok/s on Gemma 2 9B demonstrates the architecture's strength. Optimized for sparse training and inference at unprecedented scale. Not a GPU replacement — requires the CS-3 system ($2M+) and is only accessible via Cerebras cloud. Best for: organizations with frontier-scale AI budgets needing maximum throughput on supported model architectures.

Verdict

Cerebras WSE-3 (CS-3) with 44GB VRAM at 23000W TDP, scored across 4 workloads with 5 benchmark records.

Reference workload

Llama 3 8B FP16

1850 tok/s

Quantization

FP16

8K context · batch 1

LLM Inference Performance

0500100015002000Llama 370B Q4Llama 3 8BFP16Mistral 7BGemma 29B

Benchmarks (4 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

450.0tok/sFP168K, Curated Aggregate2025-03-29
Llama 3 8B FP16

llm

1850.0tok/sFP168K, Curated Aggregate2025-07-26
Mistral 7B

llm

1620.0tok/sFP168K, Curated Aggregate2024-09-12
Gemma 2 9B

llm

1420.0tok/sFP168K, Curated Aggregate2024-10-30

MyAI Score: not scored

There is insufficient comparable evidence to score Cerebras WSE-3 (CS-3). A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 44GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Good
Won't fit
70B
Tight
Won't fit
Won't fit
Reported batch-one examples
Llama 3 8B FP161850 tok/s
Llama 3 70B Q4450 tok/s
Gemma 2 9B1420 tok/s

Source

Vendor benchmark — Cerebras Inference.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2025-07-26; 410 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $2.50M only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.