The M4 Max vs RTX 4090 decision for local LLMs isn't about which is 'better', it's about which problem you're solving. If your entire workflow fits within 24GB VRAM, the RTX 4090 is absurdly more powerful, cheaper, and better supported by the CUDA ecosystem. You get faster token generation, quicker prompt processing, and the ability to scale via multi-GPU setups if needed. The power draw and heat are real, but for performance per dollar, the 4090 demolishes the M4 Max in its lane.
The M4 Max's killer argument is the unified memory capacity. No consumer GPU can match 128GB of fast memory shared between CPU and GPU, and that enables running large quantized models (70B, 120B, 140B) that are impossible on a 4090. For researchers, tinkerers, or anyone who needs local access to big models without cloud costs, the Mac Studio is the only game in town below $10k. Just know you sacrifice raw throughput, software maturity, and upgradeability. Choose based on your model size ceiling, not a spec sheet race.