How fast is Raspberry Pi 5 + AI HAT+ (Hailo-8) for local AI workloads?
Raspberry Pi 5 + AI HAT+ (Hailo-8) has a source-attributed result of 0.6 x RT on Whisper transcription (batch 1, 0-token context, INT8; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.