Llama 3 70B (Q4_K_M)
Meta Llama 3 70B Instruct, Q4_K_M GGUF, batch 1, 4K context, with offloading when VRAM-bound.
Primary metric: Tokens / sec (tok/s)
Reference prompts
Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.
- Prompt 1
A train leaves station A at 9:00 going 60 mph. Another leaves station B (180 miles away) at 9:30 going 80 mph toward A. When do they meet? Show your work.
- Prompt 2
If a 70B-parameter model uses ~140GB at FP16, what is its likely memory footprint at Q4_K_M? Show the arithmetic.
- Prompt 3
Walk through which of these is cheaper for 10M tokens/day at 30 days: H100 cloud at $2/hr vs RTX 4090 owned at $1600 + $0.10/kWh.
- Prompt 4
Explain step-by-step why FP8 inference can be 2x faster than FP16 on H100 but not on A100.
- Prompt 5
Given a 4-bit quantized Llama 3 70B model at ~40GB, how much VRAM headroom is needed for KV cache at 8K context, batch 1?
Reference runtime command
A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.
llama-server -m llama-3-70b-instruct.Q4_K_M.gguf -c 4096 -ngl 999 --seed 42Single batch. Q4_K_M (~40GB). When VRAM-bound we note `ngl` value explicitly. Median of 5 runs.
Full leaderboard
Every record for Llama 3 70B Q4.
No comparable multi-device results
The current selection lacks two devices with matching workload, quantization, context, batch and documented runtime. No speed winner is assigned. Inspect the source-attributed records.
Cite this benchmark
Use this in your paper, blog post, or comparison table.
@misc{myaihardware_llama3-70b-q4_2026,
title = {MyAI Bench: Llama 3 70B (Q4_K_M)},
author = {{MyAIHardware Contributors}},
year = {2026},
url = {https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4},
note = {Version 1.3, accessed 2026-09-09}
}MyAIHardware Contributors. (2026). MyAI Bench: Llama 3 70B (Q4_K_M). MyAIHardware. Retrieved 2026-09-09, from https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4
MyAIHardware Contributors. "MyAI Bench: Llama 3 70B (Q4_K_M)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4. Accessed 2026-09-09.
MyAI Bench, Llama 3 70B (Q4_K_M). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/llama3-70b-q4 (accessed 2026-09-09).