How fast is Intel Core Ultra 9 288V (Lunar Lake) NPU for local AI workloads?
Intel Core Ultra 9 288V (Lunar Lake) NPU has a source-attributed result of 22.0 tok/s on Phi-3 Mini (batch 1, 2048-token context, INT4; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.