Apple SiliconApple

Apple M3 Ultra (80c GPU, 512GB)

Curated Aggregate·2026-01-2610 workloads · 15 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

512 GB

TDP

270 W

MSRP

$9.5k

Perf/W

0.05 emb/s/W

Cost/1K tok

$0.0072/k

Tested

2026-01-26

Quick answer

How many tokens per second does Apple M3 Ultra (80c GPU, 512GB) produce on Llama 3 70B Q4?

Apple M3 Ultra (80c GPU, 512GB) has a source-attributed result of 14.0 tok/s on Llama 3 70B Q4 (batch 1, 8192-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-m3ultra-l3-70b-q4)As of 2026-01-26

Verdict

Apple M3 Ultra (80c GPU, 512GB) with 512GB VRAM at 270W TDP, scored across 10 workloads with 15 benchmark records.

Reference workload

Embedding throughput

2680 emb/s

Quantization

FP16

1K context · batch 1

LLM Inference Performance

020406080Llama 370B Q4Llama 3 8BQ4Llama 370B Q8Qwen 2.514BMistral 7BDeepSeek-R17B

Benchmarks (10 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

14.0tok/sQ4_K_M8K, Curated Aggregate2026-01-26
Llama 3 8B Q4

llm

38.0tok/sQ4_K_M128K, Curated Aggregate2025-03-28
Llama 3 70B Q8

llm

9.0tok/sQ8_08K, Curated Aggregate2025-04-02
Qwen 2.5 14B

llm

52.0tok/sQ4_K_M32K, Curated Aggregate2025-12-18
Mistral 7B

llm

68.0tok/sQ4_K_M16K, Curated Aggregate2025-12-04
DeepSeek-R1 7B

llm

75.0tok/sQ4_K_M32K, Curated Aggregate2026-01-08
SDXL image gen

image

4.2img/minFP16, , Curated Aggregate2026-01-22
Embedding throughput

embedding

2680.0emb/sFP161K, Curated Aggregate2026-02-08
Gemma 2 9B

llm

62.0tok/sQ4_K_M8K, Curated Aggregate2024-08-02
Asure 12B

llm

45.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score Apple M3 Ultra (80c GPU, 512GB). A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 512GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Excellent
Reported batch-one examples
Embedding throughput2680 emb/s
Llama 3 70B Q414 tok/s
SDXL image gen4 img/min

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Stale

Source-linked row with explicit verification status.

Record date: 2026-02-08; 213 days old.

Open primary source

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

Amazon is a fast way to check live availability, but memory tier and exact SKU matter more than the first visible price.

Search Amazon listings
  • +Cross-check the current ask against $9.5k and any reputable open-box or used options.
  • +Verify RAM, storage, and chip binning carefully. Small spec differences change AI usability dramatically.
  • +If you plan long 24/7 inference runs, warranty and thermals matter more than a tiny discount.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.