NPUAMD

AMD Ryzen AI Max+ 395 (Strix Halo, 96GB)

Curated Aggregate·2026-01-128 workloads · 9 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

96 GB

TDP

120 W

MSRP

$2.2k

Perf/W

0.23 tok/s/W

Cost/1K tok

$0.83/M

Tested

2026-01-12

Quick answer

How many tokens per second does AMD Ryzen AI Max+ 395 (Strix Halo, 96GB) produce on Llama 3 70B Q4?

AMD Ryzen AI Max+ 395 (Strix Halo, 96GB) has a source-attributed result of 5.0 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-strix-halo-l3-70b-q4)As of 2026-01-18

Overview

APU Memory Titan

AMD Ryzen AI Max+ 395 (Strix Halo) — flagship AI PC APU with 16 Zen 5 cores, 40 RDNA 3.5 CUs, 50 TOPS XDNA 2 NPU, and up to 96GB unified LPDDR5X memory. 120W TDP. AMD's most powerful integrated AI platform.

AI Usefulness

96GB unified memory in a laptop/desktop APU — the highest unified memory capacity in a consumer x86 chip. Runs 70B Q4 at ~30 tok/s, 32B Q4 comfortably. The XDNA 2 NPU offloads smaller models for power-efficient inference. Best for: AI laptops and compact desktops where unified memory capacity matters more than raw GPU throughput. Bridges the gap between discrete GPU AI and Apple Silicon unified memory.

Verdict

AMD Ryzen AI Max+ 395 (Strix Halo, 96GB) with 96GB VRAM at 120W TDP, scored across 8 workloads with 9 benchmark records.

Reference workload

DeepSeek-R1 7B

24 tok/s

Quantization

Q4_K_M

8K context · batch 1

LLM Inference Performance

020406080Llama 3 8BQ4Llama 370B Q4Qwen 2.514BMistral 7BPhi-3 MiniDeepSeek-R17B

Benchmarks (8 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 8B Q4

llm

28.0tok/sQ4_K_M8K, Curated Aggregate2026-01-12
Llama 3 70B Q4

llm

5.0tok/sQ4_K_M4K, Curated Aggregate2026-01-18
Qwen 2.5 14B

llm

14.0tok/sQ4_K_M8K, Curated Aggregate2026-02-04
Mistral 7B

llm

32.0tok/sQ4_K_M8K, Curated Aggregate2026-02-12
Phi-3 Mini

llm

78.0tok/sQ4_K_M4K, Curated Aggregate2026-02-20
Whisper transcription

audio

8.0x RTINT8, , Curated Aggregate2026-03-02
DeepSeek-R1 7B

llm

24.0tok/sQ4_K_M8K, Curated Aggregate2026-05-24
Gemma 2 9B

llm

30.0tok/sQ4_K_M8K, Curated Aggregate2025-02-14

MyAI Score: not scored

There is insufficient comparable evidence to score AMD Ryzen AI Max+ 395 (Strix Halo, 96GB). A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 96GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Good
Won't fit
Reported batch-one examples
DeepSeek-R1 7B24 tok/s
Whisper transcription8 x RT
Phi-3 Mini78 tok/s

Source

Fresh DeepSeek R1 run with ROCm 6.3.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Aging

Source-linked row with explicit verification status.

Record date: 2026-05-24; 108 days old.

Open primary source

Best place to start

Where to buy

Search first

Retailer we'd check first

Amazon search

For mainstream AI hardware, Amazon usually updates street pricing, seller availability, and shipping speed faster than most comparison sites.

Search Amazon listings
  • +Use $2.2k as your price anchor unless the part is clearly supply-constrained or newly launched.
  • +Check the exact cooler, board partner, or memory configuration before buying. The silicon may match, but noise and thermals do not.
  • +If the Amazon price looks inflated, wait or compare against recent street pricing rather than paying a panic premium.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.