How fast is Apple M2 Ultra (76c GPU, 192GB) for local AI workloads?
Apple M2 Ultra (76c GPU, 192GB) has a source-attributed result of 32.0 tok/s on Qwen 2.5 14B (batch 1, 8192-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.