The RTX 4090 is the undisputed king for local LLM inference, especially with large 30B-70B parameter models. Its architectural leaps, 12x larger L2 cache, faster memory, and native FP8 tensor cores, translate to real-world token generation speeds 2-3x faster than the 3090, making it feel like a generational upgrade. If you plan to build a serious home AI rig for the next 2-3 years and have the budget, the 4090 is the easy choice; the 3090 simply can't match its compute efficiency or transformer engine optimizations.
However, the RTX 3090 remains a compelling option for budget-conscious builders or those who need more than 24GB of VRAM. With NVLink, two 3090s still cost less than one 4090 and offer 48GB pooled memory, critical for fine-tuning larger models or running long-context inference. The 3090 also draws less power under sustained load for smaller models, and its raw compute is still excellent for 7B-13B parameter quantized models. Don't buy a 4090 expecting it to scale with multiple GPUs; for multi-GPU setups, the 3090's NVLink advantage can outweigh its single-card performance deficit.