How fast is NVIDIA GeForce RTX 4080 Super 16GB for local AI workloads?
NVIDIA GeForce RTX 4080 Super 16GB hits 6900.0 emb/s on Embedding throughput, its strongest benchmarked workload (batch 1, 512-token context, 16GB VRAM, 320W TDP). It has 18 records across 10 workloads in our database, with a MyAI Rating of 9.0/10. Llama 3 70B Q4 needs ~40GB VRAM, so check the VRAM column before assuming feasibility.