# MyAIHardware > Hardware specifications, memory estimates and benchmark references for people running local AI. Operated by Fredoline. Content may be AI-assisted; check each record's source, measurement conditions and evidence label rather than assuming every figure was independently tested. ## Public reading interfaces - [Agent guide](https://www.myaihardware.com/agents): read-only MCP, Markdown twins and limits. - [Canonical page catalog](https://www.myaihardware.com/agent-index.json): snapshot of public sitemap paths. - [September 9 hardware brief](https://www.myaihardware.com/ai-update-2026-09-09): source-based evaluation guidance, with no new benchmark measurements. - MCP endpoint: https://www.myaihardware.com/mcp. Append .md to supported page paths; /index.md is the homepage. Pages without server-rendered content may not have a twin. The dataset is 565 curated benchmark records across 13 workloads (Llama 3 8B/70B, DeepSeek-R1 7B, Qwen 2.5 14B, Mistral 7B, Phi-3 mini, Gemma 2 9B, Trendyol-LLM 12B, SDXL 1024, Whisper Large v3, BGE-Large embeddings) spanning a 600+ device catalog: consumer GPUs, pro GPUs, datacenter GPUs (H100, H200, B200, MI300X, MI355X), Apple Silicon Macs (Mac mini, Mac Studio, MacBook Pro), Windows AI laptops, mini-PCs, AI workstations and servers, NPUs, ASICs, and edge devices. Every row links to a citable source: llama.cpp PR threads, MLPerf submissions, Phoronix runs, vLLM CI logs, vendor whitepapers, or r/LocalLLaMA reproducible posts. Numbers are guidance, not contracts. Real-world VRAM and tokens-per-second vary with KV-cache length, batch size, quant tensor mix, and framework choice (llama.cpp, vLLM, TensorRT-LLM, MLX, ExLlamaV2). Where a single citable source does not exist, the row is flagged AGG and is the median of at least 3 independent measurements with outliers >1.5x median dropped. The full dataset is published as CC-BY-4.0 JSON and CSV at /api/v1/benchmarks.json and /api/v1/benchmarks.csv. No keys, no rate limits. An open-source CLI (myai-bench, npm) emits schema-matching JSON from llama.cpp and vLLM runs so anyone can submit reproducible numbers. ## Core reference - [Homepage](https://www.myaihardware.com/): Site overview, featured benchmarks, top buyer's guides. - [Methodology](https://www.myaihardware.com/methodology): How every number is sourced. Lists the 7 trusted source classes (llama.cpp PRs, Ollama community runs, Hugging Face hardware reports, MLPerf, Phoronix, vendor press, vLLM logs) and the aggregate-median rule for AGG rows. - [Benchmark leaderboard](https://www.myaihardware.com/benchmarks): All 565 records across 13 workloads, filterable by device class. - [Benchmark changelog](https://www.myaihardware.com/benchmarks/changelog): Every dataset version delta with timestamps. - [Public API](https://www.myaihardware.com/api/v1/benchmarks.json): Full dataset, JSON, CC-BY-4.0. - [About / E-E-A-T](https://www.myaihardware.com/about): Who runs the site, conflict-of-interest disclosure. ## Interactive tools - [LLM VRAM Calculator](https://www.myaihardware.com/llm-vram-calculator): Exact VRAM by parameter count and quant (FP16, Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q4_0, Q3_K_M, Q2_K) with the 1.10 framework-overhead constant. - [Model to GPU matcher](https://www.myaihardware.com/model-to-gpu): Pick a model, get the cheapest GPU that fits. - [What can it run](https://www.myaihardware.com/what-can-it-run): Reverse lookup, pick a GPU, see every model that fits. - [Cost-of-Inference calculator](https://www.myaihardware.com/cost-calculator): On-prem amortized cost vs API price per million tokens. - [Build wizard](https://www.myaihardware.com/build): 5-question funnel to a parts list. - [Price tracker](https://www.myaihardware.com/price-tracker): GPU street price over time. ## Top hardware reference pages - [Best GPU for local LLMs](https://www.myaihardware.com/best-gpu-for-local-llms): Ranked picks from 8 GB RTX 3060 up to 192 GB MI300X. - [Ollama hardware guide](https://www.myaihardware.com/ollama-hardware-guide): What runs on what. - [Best AI workstation](https://www.myaihardware.com/best-ai-workstation): Single-card and 4-card builds. - [DeepSeek local guide](https://www.myaihardware.com/deepseek-local-guide): R1 671B and the distill family. - [Local AI server builds](https://www.myaihardware.com/local-ai-server-builds): 24/7 rack and homelab references. - [Best mini PC for AI](https://www.myaihardware.com/best-mini-pc-for-ai): Strix Halo, Mac mini, Jetson Orin Nano Super. - [AI homelab](https://www.myaihardware.com/ai-homelab): Sub-$3K builds that run 70B at usable speeds. - [llama.cpp benchmarks](https://www.myaihardware.com/llama-cpp-benchmarks): Cross-card tokens/sec at fixed prompts. - [GPU database](https://www.myaihardware.com/gpu-database): All cards, sortable by VRAM, bandwidth, AI TOPS, MSRP. - [NPU database](https://www.myaihardware.com/npu-database): Snapdragon X, Core Ultra, Ryzen AI, Apple ANE. - [Edge database](https://www.myaihardware.com/edge-database): Jetson, Raspberry Pi + Hailo, Coral. - [ASIC database](https://www.myaihardware.com/asic-database): Groq LPU, Cerebras WSE-3, Google TPU v5p, AWS Trainium2. ## Per-model deep dives (what hardware to buy) - [Llama 3.1 8B](https://www.myaihardware.com/model/llama-3-1-8b): 8 GB at Q4, 8 GB cards or better. - [Llama 3.1 70B](https://www.myaihardware.com/model/llama-3-1-70b): 40 GB at Q4, single H100 or dual RTX 3090. - [Llama 3.3 70B](https://www.myaihardware.com/model/llama-3-3-70b): Same as 3.1 70B, newer instruction tuning. - [DeepSeek-R1 671B](https://www.myaihardware.com/model/deepseek-r1-671b): MoE, 37B active. 4x H100 80GB or 8x A100 minimum. - [DeepSeek-R1-Distill-Qwen 32B](https://www.myaihardware.com/model/deepseek-r1-distill-qwen-32b): 20 GB at Q4, fits a single 24 GB card. - [Qwen 2.5 32B](https://www.myaihardware.com/model/qwen-2-5-32b): 20 GB at Q4, the local sweet spot. - [Mixtral 8x7B](https://www.myaihardware.com/model/mixtral-8x7b): 28 GB at Q4, dual 3090 or single 48 GB pro card. - [Phi-4 14B](https://www.myaihardware.com/model/phi-4-14b): 9 GB at Q4, runs on 12 GB cards. ## Buyer's guides - [Best GPU under $500 for AI](https://www.myaihardware.com/guides/best-gpu-under-500-for-ai-2026) - [Best GPU under $1,000 for AI](https://www.myaihardware.com/guides/best-gpu-under-1000-for-ai-2026) - [Best GPU under $2,000 for AI](https://www.myaihardware.com/guides/best-gpu-under-2000-for-ai-2026) - [Best AI workstation under $3,500](https://www.myaihardware.com/guides/best-ai-workstation-3500) - [Best AI workstation under $8,000](https://www.myaihardware.com/guides/best-ai-workstation-8000) - [Best Mac for local AI](https://www.myaihardware.com/guides/best-mac-for-local-ai-2026) - [Best laptop for local LLMs](https://www.myaihardware.com/guides/best-laptop-for-local-llms-2026) - [Best AI homelab under $3,000](https://www.myaihardware.com/guides/best-ai-homelab-under-3000) - [Best used GPU for LLMs](https://www.myaihardware.com/guides/best-used-gpu-for-llms-2026) ## Head-to-head comparisons - [RTX 5090 vs RTX 4090 for AI](https://www.myaihardware.com/compare/rtx-5090-vs-rtx-4090-ai) - [M4 Max vs RTX 4090 for local LLM](https://www.myaihardware.com/compare/m4-max-vs-rtx-4090-local-llm) - [Mac Studio M3 Ultra vs dual RTX 4090](https://www.myaihardware.com/compare/mac-studio-m3-ultra-vs-dual-rtx-4090) - [AMD MI300X vs NVIDIA H100](https://www.myaihardware.com/compare/amd-mi300x-vs-nvidia-h100) - [2x RTX 3090 vs 1x RTX 4090 vs 1x RTX 5090](https://www.myaihardware.com/compare/2x-rtx-3090-vs-1x-rtx-4090-vs-1x-rtx-5090) ## Vendor hubs - [NVIDIA](https://www.myaihardware.com/vendor/nvidia) - [AMD](https://www.myaihardware.com/vendor/amd) - [Apple](https://www.myaihardware.com/vendor/apple) - [Intel](https://www.myaihardware.com/vendor/intel) - [Qualcomm](https://www.myaihardware.com/vendor/qualcomm) - [Google TPU](https://www.myaihardware.com/vendor/google-tpu) - [AWS Silicon](https://www.myaihardware.com/vendor/aws-silicon) ## Reference - [Glossary](https://www.myaihardware.com/glossary): 30 explainers covering KV cache, paged attention, quant formats, NVLink, HBM3e, MoE routing, speculative decoding. - [Deep dives](https://www.myaihardware.com/deep-dives): 20 long-form articles on CUDA vs ROCm vs Metal, FlashAttention 1/2/3, RAG architecture, tensor vs pipeline parallelism, the economics of 70B+ on-prem. - [Sitemap](https://www.myaihardware.com/sitemap.xml) - [Full corpus](https://www.myaihardware.com/llms-full.txt): The methodology and the 5 highest-value technical articles in plain text. ## Author Fredoline (contact@myaihardware.com). Independent. No NVIDIA, AMD, Apple, Intel, or vendor sponsorship. Amazon affiliate (fredoline-20) disclosed on every product page. Corrections accepted at /benchmarks/submit or via email; we update rows when a citable source contradicts them.