llm workload · 1 runs on record

Qwen3.6 35B-A3B (Q4_K_M)

Alibaba Qwen3.6 35B-A3B (35B), Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    A train leaves station A at 9:00 going 60 mph. Another leaves station B (180 miles away) at 9:30 going 80 mph toward A. When do they meet? Show your work.

  • Prompt 2

    If a 70B-parameter model uses ~140GB at FP16, what is its likely memory footprint at Q4_K_M? Show the arithmetic.

  • Prompt 3

    Walk through which of these is cheaper for 10M tokens/day at 30 days: H100 cloud at $2/hr vs RTX 4090 owned at $1600 + $0.10/kWh.

  • Prompt 4

    Explain step-by-step why FP8 inference can be 2x faster than FP16 on H100 but not on A100.

  • Prompt 5

    Given a 4-bit quantized Llama 3 70B model at ~40GB, how much VRAM headroom is needed for KV cache at 8K context, batch 1?

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
OLLAMA_NUM_PARALLEL=1 ollama run qwen-3.6-35b-a3b --verbose   # num_ctx 4096, batch 1

Alibaba Qwen3.6 35B-A3B (35B) via Ollama 0.32.1, Q4_K_M GGUF, batch 1, 4K context, measured on RTX 5090 (Vast.ai). Throughput reported as the median Ollama eval rate.

Full leaderboard

Every record for Qwen3.6 35B-A3B.

Top 10 leaderboard

Qwen3.6 35B-A3B (Q4_K_M) · sorted by tokens / sec

Bar chart: Top 1 devices ranked by Tokens / sec for Qwen3.6 35B-A3B (Q4_K_M). 1. NVIDIA GeForce RTX 5090 32GB at 169.9 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·4K ctx
8
170tok/s
32 GB575 W$2.0k0.29 tok/s/W$0.12/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_qwen-3-6-35b-a3b-q4_2026,
  title  = {MyAI Bench: Qwen3.6 35B-A3B (Q4_K_M)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/qwen-3-6-35b-a3b-q4},
  note   = {Version 1.3, accessed 2026-08-30}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Qwen3.6 35B-A3B (Q4_K_M). MyAIHardware. Retrieved 2026-08-30, from https://www.myaihardware.com/benchmarks/workload/qwen-3-6-35b-a3b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Qwen3.6 35B-A3B (Q4_K_M)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/qwen-3-6-35b-a3b-q4. Accessed 2026-08-30.
Plain text
MyAI Bench, Qwen3.6 35B-A3B (Q4_K_M). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/qwen-3-6-35b-a3b-q4 (accessed 2026-08-30).