ASICGoogle

Google TPU v5e

Curated Aggregate·2024-08-084 workloads · 4 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

16 GB

TDP

170 W

MSRP

$9.0k

Perf/W

1.41 tok/s/W

Cost/1K tok

$0.40/M

Tested

2024-08-08

Quick answer

How fast is Google TPU v5e for local AI workloads?

Google TPU v5e has a source-attributed result of 320.0 tok/s on Phi-3 Mini (batch 1, 4096-token context, INT8; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-x3-tpu-v5e-phi3)As of 2024-08-26

Verdict

Google TPU v5e with 16GB VRAM at 170W TDP, scored across 4 workloads with 4 benchmark records.

Reference workload

Phi-3 Mini

320 tok/s

Quantization

INT8

4K context · batch 1

LLM Inference Performance

080160240320Llama 3 8BFP16Llama 370B Q8Mistral 7BPhi-3 Mini

Benchmarks (4 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B FP16

llm

240.0tok/sFP164K, Curated Aggregate2024-08-08
Llama 3 70B Q8

llm

62.0tok/sINT84K, Curated Aggregate2024-08-15
Mistral 7B

llm

195.0tok/sINT84K, Curated Aggregate2024-06-15
Phi-3 Mini

llm

320.0tok/sINT84K, Curated Aggregate2024-08-26

MyAI Score: not scored

There is insufficient comparable evidence to score Google TPU v5e. A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 16GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Won't fit
13-14B
Excellent
Tight
Won't fit
32B
Won't fit
Won't fit
Won't fit
70B
Won't fit
Won't fit
Won't fit
Reported batch-one examples
Phi-3 Mini320 tok/s
Llama 3 70B Q862 tok/s
Llama 3 8B FP16240 tok/s

Source

Phi-3 Mini on TPU v5e.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2024-08-26; 744 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $9.0k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.