How many tokens per second does AMD Threadripper PRO 7995WX (96-core) produce on Llama 3 70B Q4?
AMD Threadripper PRO 7995WX (96-core) has a source-attributed result of 2.4 tok/s on Llama 3 70B Q4 (batch 1, 4096-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.