llm workload · 1 runs on record

Granite 4.1 30B (Q4_K_M)

IBM Granite 4.1 30B (30B), Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Explain the relationship between attention heads and KV cache memory usage for a 7B-parameter transformer at 4K context.

  • Prompt 2

    Write a 60-word product description for a mid-range AI workstation: 1× RTX 4090, 64GB DDR5, 2TB NVMe. Highlight one trade-off.

  • Prompt 3

    List five concrete differences between INT4 weight-only quantization (Q4_K_M) and INT8 quantization (Q8_0) for inference.

  • Prompt 4

    I have 12GB of VRAM and want to run a coding assistant locally. What model size and quantization fit, with what context length?

  • Prompt 5

    Summarize the difference between greedy decoding and nucleus sampling in 3 short bullet points.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
OLLAMA_NUM_PARALLEL=1 ollama run granite-4.1-30b --verbose   # num_ctx 4096, batch 1

IBM Granite 4.1 30B (30B) via Ollama 0.32.1, Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090 (Vast.ai). Throughput reported as the median Ollama eval rate.

Full leaderboard

Every record for Granite 4.1 30B.

Top 10 leaderboard

Granite 4.1 30B (Q4_K_M) · sorted by tokens / sec

Bar chart: Top 1 devices ranked by Tokens / sec for Granite 4.1 30B (Q4_K_M). 1. NVIDIA GeForce RTX 5090 32GB at 73.3 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
8
73tok/s
32 GB575 W$2.0k0.13 tok/s/W$0.29/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_granite-4-1-30b-q4_2026,
  title  = {MyAI Bench: Granite 4.1 30B (Q4_K_M)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/granite-4-1-30b-q4},
  note   = {Version 1.3, accessed 2026-08-30}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Granite 4.1 30B (Q4_K_M). MyAIHardware. Retrieved 2026-08-30, from https://www.myaihardware.com/benchmarks/workload/granite-4-1-30b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Granite 4.1 30B (Q4_K_M)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/granite-4-1-30b-q4. Accessed 2026-08-30.
Plain text
MyAI Bench, Granite 4.1 30B (Q4_K_M). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/granite-4-1-30b-q4 (accessed 2026-08-30).