llm workload · 21 runs on record

Trendyol-LLM-Asure 12B (Q4)

alibayram/Trendyol-LLM-Asure-12B, Q4_K_M GGUF, batch 1, 8K context. Turkish-optimized LLM benchmark.

Primary metric: Tokens / sec (tok/s)

Reference prompts

Representative prompts for this workload. Exact prompts and harness settings still depend on the cited source for each record.

  • Prompt 1

    Write a comprehensive explanation of transformer attention mechanisms in Turkish.

  • Prompt 2

    Solve this coding problem: implement a binary search tree in Python with insert, delete, and search methods.

  • Prompt 3

    Analyze the following Turkish legal text and summarize the key points in English.

Reference runtime command

A representative invocation for reproducing this workload class. Source-specific runs may use adjacent runtimes unless the record says otherwise.

shell
llama-cli -m Trendyol-LLM-Asure-12B-Q4_K_M.gguf -p PROMPT -n 512 -c 8192 -ngl 99

alibayram/Trendyol-LLM-Asure-12B via llama.cpp. Q4_K_M quant, batch 1, 8K context. Test bilingual (TR+EN) reasoning, coding, and summarization capabilities.

Full leaderboard

Every record for Asure 12B.

Top 10 leaderboard

Trendyol-LLM-Asure 12B (Q4) · sorted by tokens / sec

Bar chart: Top 10 devices ranked by Tokens / sec for Trendyol-LLM-Asure 12B (Q4). 1. Groq LPU Inference Engine at 520 tok/s. 2. NVIDIA B200 192GB at 350 tok/s. 3. NVIDIA H200 141GB at 230 tok/s. 4. NVIDIA H100 SXM5 80GB at 195 tok/s. 5. NVIDIA A100 SXM4 80GB at 145 tok/s.

Value frontier, MSRP vs tokens / sec

Each dot is a device. Top-left is best value (cheap + fast).

Scatter chart of MSRP versus Tokens / sec across 21 devices. Devices in the upper-left region offer the best price/performance ratio. Top 5 by primary metric: Groq LPU Inference Engine at $20,000 delivering 520 tok/s; B200 192GB at $39,999 delivering 350 tok/s; H200 141GB at $30,000 delivering 230 tok/s; H100 SXM5 80GB at $25,000 delivering 195 tok/s; A100 SXM4 80GB at $15,000 delivering 145 tok/s.

Full leaderboard

Click any row for detailed breakdown. Click column headers to sort.

#DeviceVerifBuy
1
Groq LPU Inference EngineASIC
Groq·FP8·8K ctx
5
520tok/s
σ 15.6
230 GB215 W$20k2.42 tok/s/W$0.41/MAmazon
2
NVIDIA B200 192GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
350tok/s
σ 10.5
192 GB1000 W$40k0.35 tok/s/W$0.0012/kAmazon
3
NVIDIA H200 141GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
230tok/s
σ 9.2
141 GB700 W$30k0.33 tok/s/W$0.0014/kAmazon
#4
NVIDIA H100 SXM5 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
195tok/s
σ 7.8
80 GB700 W$25k0.28 tok/s/W$0.0014/kAmazon
#5
NVIDIA A100 SXM4 80GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
145tok/s
σ 5.8
80 GB400 W$15k0.36 tok/s/W$0.0011/kAmazon
#6
AMD Instinct MI300X 192GBDatacenter GPU
AMD·Q4_K_M·8K ctx
6
130tok/s
σ 6.5
192 GB750 W$15k0.17 tok/s/W$0.0012/kAmazon
#7
NVIDIA L40S 48GBDatacenter GPU
NVIDIA·Q4_K_M·8K ctx
6
110tok/s
σ 5.5
48 GB350 W$7.8k0.31 tok/s/W$0.75/MAmazon
#8
NVIDIA DGX Spark (Project DIGITS, 128GB)Datacenter GPU
NVIDIA·Q4_K_M·8K ctx
5
100tok/s
σ 5.0
128 GB240 W$3.0k0.42 tok/s/W$0.32/MAmazon
#9
NVIDIA RTX 6000 Ada 48GBPro GPU
NVIDIA·Q4_K_M·8K ctx
6
95tok/s
σ 4.8
48 GB300 W$6.8k0.32 tok/s/W$0.76/MAmazon
#10
NVIDIA GeForce RTX 5090 32GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
85tok/s
σ 4.2
CI: 80.3–89.8
32 GB575 W$2.0k0.15 tok/s/W$0.25/MAmazon
#11
NVIDIA GeForce RTX 5080 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
10
70tok/s
σ 1.5
CI: 68.6–71.2
16 GB360 W$9990.19 tok/s/W$0.15/MAmazon
#12
NVIDIA GeForce RTX 5070 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
63tok/s
σ 3.1
CI: 59.5–66.5
16 GB300 W$7490.21 tok/s/W$0.13/MAmazon
#13
NVIDIA GeForce RTX 4090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
60tok/s
σ 3.2
CI: 56.4–63.6
24 GB450 W$1.6k0.13 tok/s/W$0.28/MAmazon
#14
NVIDIA GeForce RTX 4080 Super 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
59tok/s
σ 3.0
CI: 55.6–62.4
16 GB320 W$9990.18 tok/s/W$0.18/MAmazon
#15
AMD Radeon RX 7900 XTX 24GBConsumer GPU
AMD·Q4_K_M·8K ctx
6
52tok/s
σ 3.6
24 GB355 W$9990.15 tok/s/W$0.20/MAmazon
#16
Apple M3 Ultra (80c GPU, 512GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
45tok/s
σ 2.7
512 GB100 W$10.0k0.45 tok/s/W$0.0023/kAmazon
#17
NVIDIA GeForce RTX 4070 Ti 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
44tok/s
σ 2.5
CI: 41.2–46.8
12 GB285 W$7990.15 tok/s/W$0.19/MAmazon
#18
NVIDIA GeForce RTX 3090 24GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
38tok/s
σ 2.3
CI: 35.4–40.6
24 GB350 W$1.5k0.11 tok/s/W$0.42/MAmazon
#19
Apple M4 Max (40c GPU, 128GB)Apple Silicon
Apple·Q4_K_M·8K ctx
6
35tok/s
σ 2.5
128 GB65 W$4.7k0.54 tok/s/W$0.0014/kAmazon
#20
NVIDIA GeForce RTX 4060 Ti 16GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
28tok/s
σ 1.7
CI: 26.1–29.9
16 GB165 W$4990.17 tok/s/W$0.19/MAmazon
#21
NVIDIA GeForce RTX 3060 12GBConsumer GPU
NVIDIA·Q4_K_M·8K ctx
6
20tok/s
σ 1.5
CI: 18.3–21.7
12 GB170 W$3290.12 tok/s/W$0.17/MAmazon
Sorted by Tokens / sec (high → low)

Cite this benchmark

Use this in your paper, blog post, or comparison table.

BibTeX
@misc{myaihardware_trendyol-llm-asure-12b-q4_2026,
  title  = {MyAI Bench: Trendyol-LLM-Asure 12B (Q4)},
  author = {{MyAIHardware Contributors}},
  year   = {2026},
  url    = {https://www.myaihardware.com/benchmarks/workload/trendyol-llm-asure-12b-q4},
  note   = {Version 1.3, accessed 2026-08-27}
}
APA
MyAIHardware Contributors. (2026). MyAI Bench: Trendyol-LLM-Asure 12B (Q4). MyAIHardware. Retrieved 2026-08-27, from https://www.myaihardware.com/benchmarks/workload/trendyol-llm-asure-12b-q4
MLA
MyAIHardware Contributors. "MyAI Bench: Trendyol-LLM-Asure 12B (Q4)." MyAIHardware, 2026, https://www.myaihardware.com/benchmarks/workload/trendyol-llm-asure-12b-q4. Accessed 2026-08-27.
Plain text
MyAI Bench, Trendyol-LLM-Asure 12B (Q4). MyAIHardware Contributors, 2026. Version 1.3. https://www.myaihardware.com/benchmarks/workload/trendyol-llm-asure-12b-q4 (accessed 2026-08-27).