Local AI Server Builds: From Dev Box to Rackmount
Unvalidated planning example. The amounts are undated estimates from an earlier parts list, not merchant observations. Cooling, tax, delivery and any missing adapters are excluded. Obtain a compatible, itemized quote before ordering.
Reference subtotals at a glance
12GB server concept
12GB GPU concept for a small headless inference service. Throughput, noise and concurrent-user capacity are not established for the whole system. Confirm GPU width, RAM support and cooling.
Dual-3090 server concept
Dual RTX 3090 concept with 48GB total GPU memory. Ryzen/X570 uses unbuffered DIMMs; the former registered-DIMM selection is incompatible and withdrawn. Validate a UDIMM kit, PCIe layout and a supported model-splitting runtime before specifying a system.
Four-GPU EPYC concept
Four RTX 4090 concept with 96GB total GPU memory. The chassis, GPU spacing, airflow and power configuration have not been validated. A qualified integrator must establish a compatible configuration before purchase.
Rack server concept — pairing withdrawn
The earlier AS-4124GS-TNR / EPYC 9554 / DDR5 combination is withdrawn: the named 4U server belongs to an older platform and is not a verified match for those components. Do not order this combination. Request a compatible current platform and a complete quote.
12GB server concept
Undated line-item estimates · purchase list unvalidated
12GB GPU concept for a small headless inference service. Throughput, noise and concurrent-user capacity are not established for the whole system. Confirm GPU width, RAM support and cooling.
Wins
- Capacity target only; no whole-system performance ranking
Drawbacks
- Incomplete quote; cooling and ancillary costs excluded
- Purchase only after platform compatibility is confirmed
CPU
AMD Ryzen 7 7700 (8C/16T)
Undated line-item estimate; exact SKU and compatibility need confirmation.
GPU
NVIDIA RTX 4070 Super 12 GB
Undated line-item estimate; exact SKU and compatibility need confirmation.
RAM
64GB DDR5 UDIMM (2×32); exact kit pending
Do not assume ECC support on this board; confirm the CPU, BIOS and memory QVL. This amount is a prior allowance.
Storage
Samsung 990 Pro 4 TB NVMe + 8 TB Seagate IronWolf
Undated line-item estimate; exact SKU and compatibility need confirmation.
Motherboard
Gigabyte B650M Aorus Elite AX
Undated line-item estimate; exact SKU and compatibility need confirmation.
PSU
Corsair RM850e (ATX 3.1)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Case
Fractal Design Define 7 Compact
Undated line-item estimate; exact SKU and compatibility need confirmation.
Dual-3090 server concept
Undated line-item estimates · purchase list unvalidated
Dual RTX 3090 concept with 48GB total GPU memory. Ryzen/X570 uses unbuffered DIMMs; the former registered-DIMM selection is incompatible and withdrawn. Validate a UDIMM kit, PCIe layout and a supported model-splitting runtime before specifying a system.
Wins
- Capacity target only; no whole-system performance ranking
Drawbacks
- Incomplete quote; cooling and ancillary costs excluded
- Purchase only after platform compatibility is confirmed
CPU
AMD Ryzen 9 5950X (16C/32T)
Undated line-item estimate; exact SKU and compatibility need confirmation.
GPU ×2
2× NVIDIA RTX 3090 24 GB + NVLink bridge
Undated line-item estimate; exact SKU and compatibility need confirmation.
RAM
128GB DDR4 UDIMM (4×32); compatible kit not selected
RDIMM is incompatible with the Ryzen/X570 platform. The amount shown is the previous budget allowance, not a quote for a replacement kit.
Storage
Crucial T705 2 TB + 4× 8 TB WD Red Plus RAID-Z2
4 × 8TB in RAID-Z2 provides about 16TB before formatting, plus the separate 2TB SSD.
Motherboard
ASUS ProArt X570-Creator WiFi
Verify x8/x8 slot wiring, BIOS and physical GPU spacing in the manual; no PLX switch is established.
PSU
Seasonic PRIME PX-1300 ATX 3.1
Undated line-item estimate; exact SKU and compatibility need confirmation.
Case
Fractal Define 7 XL
Undated line-item estimate; exact SKU and compatibility need confirmation.
Network
MikroTik CRS305-1G-4S+IN 10 GbE switch
Undated line-item estimate; exact SKU and compatibility need confirmation.
Four-GPU EPYC concept
Undated line-item estimates · purchase list unvalidated
Four RTX 4090 concept with 96GB total GPU memory. The chassis, GPU spacing, airflow and power configuration have not been validated. A qualified integrator must establish a compatible configuration before purchase.
Wins
- Capacity target only; no whole-system performance ranking
Drawbacks
- Incomplete quote; cooling and ancillary costs excluded
- Purchase only after platform compatibility is confirmed
CPU
AMD EPYC 9354 (32C/64T)
Undated line-item estimate; exact SKU and compatibility need confirmation.
GPU ×4
4× NVIDIA RTX 4090 24 GB
Undated line-item estimate; exact SKU and compatibility need confirmation.
RAM
256 GB DDR5-4800 ECC RDIMM (8×32)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Storage
Solidigm D7-PS1010 3.84 TB U.2 ×2 (mirror)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Motherboard
Asus K14PA-U24 EPYC server board
Undated line-item estimate; exact SKU and compatibility need confirmation.
PSU
FSP CUP-T2 2000 W Platinum
Undated line-item estimate; exact SKU and compatibility need confirmation.
Case
SilverStone RM43-320-RS 4U rackmount
Undated line-item estimate; exact SKU and compatibility need confirmation.
Network
Mellanox ConnectX-4 25 GbE
Undated line-item estimate; exact SKU and compatibility need confirmation.
Rack server concept — pairing withdrawn
Undated line-item estimates · purchase list unvalidated
The earlier AS-4124GS-TNR / EPYC 9554 / DDR5 combination is withdrawn: the named 4U server belongs to an older platform and is not a verified match for those components. Do not order this combination. Request a compatible current platform and a complete quote.
Wins
- Capacity target only; no whole-system performance ranking
Drawbacks
- Incomplete quote; cooling and ancillary costs excluded
- Purchase only after platform compatibility is confirmed
CPU ×2
2× AMD EPYC 9554 (64C/128T each)
Undated line-item estimate; exact SKU and compatibility need confirmation.
GPU ×4
4× NVIDIA RTX 6000 Ada 48 GB
Undated line-item estimate; exact SKU and compatibility need confirmation.
RAM
1 TB DDR5-4800 ECC RDIMM (16×64)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Storage
4× Solidigm D7-PS1010 7.68 TB (RAID-Z1)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Server chassis
Withdrawn pairing: Supermicro AS-4124GS-TNR
4U older-platform system; do not combine with the listed EPYC 9554 and DDR5. The amount is retained only to account for the earlier estimate.
PSU
Built-in 2× 2200 W Platinum (redundant)
Undated line-item estimate; exact SKU and compatibility need confirmation.
Networking
2× ConnectX-6 100 GbE + UniFi PRO 24 PoE
Undated line-item estimate; exact SKU and compatibility need confirmation.
Rack & UPS
StarTech 25U + APC SRT3000RMXLA 3 kVA UPS
Undated line-item estimate; exact SKU and compatibility need confirmation.
Headless Linux setup, the AI server stack
We standardize on Ubuntu Server LTS 24.04 + Docker Compose + Tailscale + Caddy. The boring choices are the right ones for an always-on server.
Base OS
# Ubuntu Server 24.04 LTS
# Pick: minimal install, OpenSSH, no GUI
sudo apt update && sudo apt full-upgrade -y
sudo apt install -y curl tmux htop nvtop \
docker.io docker-compose-v2 \
fail2ban ufwNVIDIA stack
# Driver + Container Toolkit
sudo ubuntu-drivers autoinstall
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure \
--runtime=docker
sudo systemctl restart dockerTailscale (remote access)
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up --ssh \
--advertise-tags=tag:server \
--accept-routes
# Now reachable at 100.x.x.x from anywhereDocker Compose, Ollama + WebUI
services:
ollama:
image: ollama/ollama:latest
runtime: nvidia
volumes: ["./models:/root/.ollama"]
restart: unless-stopped
webui:
image: ghcr.io/open-webui/open-webui:main
ports: ["3000:8080"]
depends_on: [ollama]Power, networking, IPMI, the un-fun parts
Power planning
A standard 15 A residential circuit safely sustains ~1,440 W (15 A × 120 V × 0.8 derating). Anything past that wants its own 20 A or 30 A circuit. Plan UPS sizing around the peak, not idle, a 1500 VA UPS covers 4× RTX 4090 during a brief brownout but cannot sustain it.
Networking
1 GbE works for serving chat. NAS-backed training datasets want 10 GbE minimum; multi-node inference (vLLM with TP>1) wants 25/100 GbE. MikroTik CRS305 (4× SFP+) and CRS504 (4× QSFP28) are the cheap-but-real switches to evaluate against your network requirements.
IPMI / BMC
Worth it on EPYC / Xeon W boards. Remote power, virtual KVM, sensor logging, virtual media for OS install. Never expose IPMI to the internet, even with a strong password. Put it on a tagged VLAN or behind Tailscale.
Cooling & noise
Multi-GPU servers run hot. Blower-style cards (RTX 6000 Ada, RTX A6000) exhaust heat out the back rather than dumping it on the next card. For consumer cards (4090, 5090), space them with at least one slot of breathing room, three-slot cards in adjacent slots will thermal-throttle.
Frequently asked questions
Can I host a local AI server in my closet?
Yes, but plan for heat and noise. A single-GPU dev box draws 350–500 W and is closet-tolerable with passive ventilation. A 4-GPU workstation draws 2,000+ W and needs active extraction or a dedicated room.
Should I run Linux or Windows on an AI server?
Linux. Driver maturity, CUDA stack stability, Docker support, and remote management (SSH, tmux, htop) are all dramatically better. Ubuntu Server LTS 24.04 or Debian 12 are the default picks.
Do I need IPMI for a home AI server?
Not strictly, Tailscale + ssh + a Pi-KVM gets you 90% there for free. IPMI is worth it on EPYC / Xeon W server boards because remote console + power control + sensor monitoring + virtual-media are bundled.
How loud is a 4-GPU AI server?
Idle, 35–45 dBA with mid-range Noctua fans. Loaded with blower-style GPUs at 80% fan, 55–65 dBA, comparable to a vacuum cleaner. Move it out of your bedroom.
Is dual EPYC overkill for AI inference?
Single EPYC is enough for almost all inference workloads, 4 GPUs hang off 64 PCIe lanes comfortably. Dual EPYC only makes sense for very high-batch inference servers or when CPU-side preprocessing is the bottleneck.
Can I expose my AI server to the internet?
Don't, at least not directly. Put it behind Tailscale, WireGuard, or a Cloudflare Tunnel. Add basic-auth to Open WebUI. A bare-internet Ollama port is asking to be turned into someone else's free GPT.
Related guides
Best AI Workstation Builds in 2026 (Every Budget)
Four AI workstation planning examples with calculated reference subtotals, memory requirements and explicit compatibility checks. Purchase lists remain unvalidated.
Read guide HomelabAI Homelab Setup: The 2026 Build Guide
Build a complete AI homelab: networking, NAS for datasets, GPU passthrough on Proxmox, Ollama + Open WebUI, local agents, and power budgeting.
Read guide GPUBest GPU for Local LLMs (2026)
GPU choices for local LLMs based on memory capacity, software support and source-attributed benchmark records. Includes limitations and reference pricing.
Read guide SoftwareRun DeepSeek Locally, R1, V3 & Coder Guide
Memory planning for DeepSeek-R1, V3 and Coder, with source-attributed evidence and explicit limits on hardware recommendations.
Read guideStay Ahead of the AI Curve
Get weekly AI hardware news, benchmark updates, and deals in your inbox. Founding-subscriber list, be one of the first.
Affiliate disclosure: As an Amazon Associate, MyAIHardware.com earns from qualifying purchases. Server builds are spec'd from neutral evaluation.