Latest cut: v1.4 · 565 source-attributed records · latest run August 2026

Local AI Benchmarks

The local-AI hardware index.

Aggregated numbers from real silicon, Llama, Mistral, DeepSeek, Qwen, SDXL, Whisper. Sortable, filterable, with $/tok·s⁻¹ and perf/watt so you can decide what to actually buy.

Where these numbers come from: These benchmarks are aggregated from llama.cpp community runs, vendor-published numbers, vLLM logs, and MLPerf submissions. MyAIHardware does not yet operate an in-house benchmark lab, that's coming Q3 2026. See per-row source notes and our full methodology.

565
Benchmark records
30
Workloads tracked
78
Devices tracked
Aug 2026
Latest data
462
Single-stream records

Pick a workload

Filter by silicon class, quantization, price, and VRAM to find your perfect setup.

Rankings show the largest available group with matching workload, quantization, context, batch-one settings and recorded runtime. Other configurations remain in the JSON and CSV exports.

Workload

Device class

Quantization

Price tier

Min VRAM

Showing 0 of 0 results

No comparable multi-device results

The current selection lacks two devices with matching workload, quantization, context, batch and documented runtime. No speed winner is assigned. Inspect the source-attributed records.

All 30 workloads

Quick reference for every benchmark we run, including what it actually measures.

External model quality layer

We separate hardware speed from model quality. These upstream snapshots show what the broader model-eval world thinks is strong right now, while our local benchmark data shows what actually runs well on your machine.

Snapshot generated

May 28, 2026

Weekly automated ingest with source-level freshness metadata.

Open-weight shortlist

6 local-friendly picks

Filtered to open text models with published parameter sizes and plausible local fit.

Reference sources

3 feeds

OpenEvals, LM Arena community Elo, and LiveBench model judgments.

Best open models to run locally

OpenEvals

Full snapshot105/105

Directional quality signal for open-weight models. We bias this slice toward text models with known parameter sizes so builders can map quality to hardware fit.

#1

microsoft/Phi-3-medium-4k-instruct

14B · est. 8.4 GB Q4

91.0 score

#2

Qwen/Qwen2-72B

73B · est. 43.6 GB Q4

89.5 score

#3

microsoft/Phi-3.5-mini-instruct

3.8B · est. 2.3 GB Q4

86.2 score

#4

internlm/internlm2_5-7b-chat

7.7B · est. 4.6 GB Q4

86.0 score

#5

microsoft/Phi-3-mini-4k-instruct

3.8B · est. 2.3 GB Q4

85.7 score

Community preference ceiling

LM Arena

Partial snapshot1K/8.9K

Useful for the broad popularity and preference frontier. This is mostly hosted frontier-model context, not a local-fit recommender.

#1

claude-opus-4-6-thinking

anthropic · 27.5K votes

#1 · 1500.0

#2

claude-opus-4-6

anthropic · 29.2K votes

#2 · 1497.9

#3

gemini-3.5-flash

google · 5.9K votes

#3 · 1486.0

#4

claude-opus-4-7-thinking

anthropic · 12.9K votes

#4 · 1485.9

#5

gemini-3.1-pro-preview

google · 34.2K votes

#5 · 1482.8

Recent task-judged performance

LiveBench

Partial snapshot2K/60.4K

Another external frontier reference. We surface it so readers can separate 'runs fast locally' from 'is judged strong on recent tasks'.

#1

claude-3-7-sonnet-20250219-base

1 task buckets · 29 evals

82.8 / 100

#2

claude-3-opus-20240229

1 task buckets · 29 evals

82.8 / 100

#3

o1-2024-12-17-high

1 task buckets · 29 evals

82.8 / 100

#4

claude-3-7-sonnet-20250219-thinking-64k

1 task buckets · 29 evals

79.3 / 100

#5

gemini-2.0-flash-thinking-exp-1219

1 task buckets · 29 evals

79.3 / 100
Methodology

How MyAI Bench measures.

Every record ships with the workload, model, quantization, runtime notes, source, and test date we could verify from the cited run. Because this is a blended dataset rather than one locked lab harness, compare results within the same workload and verification tier.

Records in database565
Unique devices78
Distinct workloads30
Distinct test dates325
Bench versionv1.4
Full bench charter

Curated sources today

This release combines cited public runs from llama.cpp, MLPerf, vendor disclosures, and community logs. We normalize the metadata and clearly tag verification status rather than pretending every row came from one in-house harness.

Reference prompts and commands

Each workload page shows a representative prompt set and runtime command so you can reproduce the class of test. Exact prompts, runtimes, and harness settings still depend on the cited source for each record.

Transparency on offloading

When a model doesn't fit in VRAM we explicitly note partial CPU offload. Pure-GPU runs and offload runs are not combined in the same rank.

Open dataset, versioned

Schema lives in src/data/benchmarkDatabase.ts on GitHub. Every record carries a sourceNote and testedAt date. v1.4 is the current cut, and the downloadable API is generated from the same enriched dataset the app renders.

Vendor results clearly tagged

Vendor-published numbers keep their sourceNote and verification label so you can apply your own discount. We do not silently blur vendor claims, community runs, and lab-verified records into one confidence tier.

Derived metrics, no fudging

$/1k tok is derived from MSRP and observed throughput using the public formula shown here. Perf/W is tokens/sec divided by TDP. No hidden weights, no proprietary composite score.

Crowdsourced data

Got numbers? Submit your bench.

Running an exotic setup, Strix Halo, 4×3090 in a frame, M3 Ultra at 512GB unified? We want the data. Submissions are reviewed and credited. Repeat contributors get early access to upcoming benchmark releases.

Llama.cpp commit + flags
Single batch, median of 5 runs
Power measured at the wall
Screenshots or log links accepted

Never Miss a Benchmark

Weekly AI hardware news + deals

Dig deeper

Cross-reference benchmarks with our hardware databases.