Hermes 3 Llama 3.1 8B (Q4_K_M)
NousResearch Hermes 3 Llama 3.1 8B (8B), Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090.
Primary metric: Tokens / sec (tok/s)
Reference prompts
Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.
- Prompt 1
Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.
- Prompt 2
Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.
- Prompt 3
List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.
- Prompt 4
I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?
- Prompt 5
Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.
Reference runtime command
A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.
OLLAMA_NUM_PARALLEL=1 ollama run hermes-3-llama-3.1-8b --verbose # num_ctx 4096, batch 1NousResearch Hermes 3 Llama 3.1 8B (8B) via Ollama 0.32.1, Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090 (Vast.ai). Throughput reported as the median Ollama eval rate.
Full leaderboard
Every record for Hermes 3 Llama 3.1 8B.
Top 10 leaderboard
Hermes 3 Llama 3.1 8B (Q4_K_M) · sorted by tokens / sec
Full leaderboard
Click any row for detailed breakdown. Click column headers to sort.
| # | Device | Verif | Buy | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | NVIDIA GeForce RTX 5090 32GBConsumer GPU NVIDIA·Q4_K_M·4K ctx | 8 | 205tok/s | 32 GB | 575 W | $2.0k | 0.36 tok/s/W | $0.10/M | Amazon |
Cite this benchmark
Use this in your paper, blog post, or comparison table.
@misc{myaihardware_hermes-3-llama-3-1-8b-q4_2026,
title = {MyAI Bench: Hermes 3 Llama 3.1 8B (Q4_K_M)},
author = {{MyAIHardware Contributors}},
year = {2026},
url = {https://www.myaihardware.com/benchmarks/workload/hermes-3-llama-3-1-8b-q4},
note = {Version 1.3, accessed 2026-08-30}
}MyAIHardware Contributors. (2026). MyAI Bench: Hermes 3 Llama 3.1 8B (Q4_K_M). MyAIHardware. Retrieved 2026-08-30, from https://www.myaihardware.com/benchmarks/workload/hermes-3-llama-3-1-8b-q4
MyAIHardware Contributors. "MyAI Bench: Hermes 3 Llama 3.1 8B (Q4_K_M)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/hermes-3-llama-3-1-8b-q4. Accessed 2026-08-30.
MyAI Bench, Hermes 3 Llama 3.1 8B (Q4_K_M). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/hermes-3-llama-3-1-8b-q4 (accessed 2026-08-30).