Observed Signal · Jun 11, 2026 · Benchmark/Performance Test · Source: DEV Community · Impact: 2/5 · Sentiment: Positive

MTP (Speculative Decoding) Is Hardware-Dependent

Executive Signal Summary

A developer benchmark shows speculative decoding (MTP) can nearly double token decode throughput on a single RTX 3090 with Gemma 4 12B QAT (1.95× at n-max 3), but the same MTP draft slows the same model on an M1 Max (0.87×). Cross-hardware results (RTX 3090, RTX 5070 Ti laptop, M1 Max) indicate MTP's benefit depends on the draft-cost-to-verify-cost ratio for a given architecture: capable CUDA GPUs with spare compute see large throughput gains, while Apple Silicon’s memory/compute balance can make MTP net overhead. The post includes reproducible settings (llama.cpp commit, model files, llama-server and speed_bench commands) and notes the 3090 runs were stable (CV < 0.5%).

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Provides practical, reproducible evidence that speculative decoding (MTP) yields large GPU speedups on capable CUDA hardware but can degrade performance on Apple Silicon—important to engineers optimizing LLM inference and cost but not industry-shifting.

SIGNAL RADAR

Track Apple Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Gemma 4 12B QAT (UD-Q4_K_XL) with a Q8_0-MTP draft achieved mean 167.4 tok/s (1.95× speedup) on a single RTX 3090 at n-max 3.
  • MTP n-max 2 on the same 3090 produced 159.4 tok/s (1.86× speedup); draft acceptance rates were ~0.77 (n-max 2) and ~0.69 (n-max 3).
  • A cross-hardware comparison reported RTX 5070 Ti at ~1.74× speedup and an M1 Max at ~0.87× (slower with MTP enabled).
  • The benchmark fit in approximately 8 GB of VRAM and the 3090 runs showed low run-to-run variance (CV < 0.5%).
  • Author provides reproduction details: llama.cpp commit e3471b3, model filenames from unsloth/gemma-4-12B-it-qat-GGUF, and sample llama-server and speed_bench.py commands.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: Jun 11, 2026
Original Coverage Title: “MTP Isn't Always a Win: 1.95x on My 3090, but Speculative Decoding Is Hardware-Dependent”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIAug 29, 2026

Gemma 4 in Pure JAX: TPU-to-GPU Port Lessons

A developer ported a Gemma 4 E2B checkpoint to a single pure-JAX codebase and ran it on Cloud TPU v5e/v6e and an NVIDIA T4G (on an AWS Graviton2 host). Most model code, compilation cache behavior, and static-shape discipline transferred unchanged, but two hardware-dependent issues surfaced: a fused W4A16 kernel written in Pallas (tiled for TPU VMEM) cannot run on GPUs due to much smaller shared-memory limits, and compute-dtype selection must be detected at runtime (float16 vs bfloat16) to avoid hidden conversion costs. The article documents Gemma 4’s four model irregularities, a KV-ring padding bug that produced silent token loops, measured decode throughput (13.10 tok/s on T4G), and profiling that shows unexpected conversion overhead on a Turing GPU.

Read assessment
Large Language Models (LLM) & AIAug 10, 2026

TileRT Enables Ultra-High Interactivity on NVIDIA GPUs

SemiAnalysis reports on TileRT, a persistent-engine approach that compiles an entire decode graph into a single resident kernel on NVIDIA GPUs to reduce per-token latency. In InferenceX benchmarks TileRT reached up to 500 tokens/s/user on a single B200 decode server (GLM5 FP8 744B) and delivered large gains versus traditional GPU inference engines (e.g., ~3× vs GB300 NVL72 in some tests). TileRT is designed to handle latency-sensitive decode while remaining interoperable with throughput-optimized prefill engines such as vLLM. TileRT is already deployed in production at Xiaomi and Z.ai, but currently supports a small model catalog (GLM-5/5.1, DeepSeek-V3.2) and primarily serves batch size 1 decode workloads, reflecting trade-offs between per-user interactivity and aggregate throughput.

Read assessment
Large Language Models (LLM) & AIMay 24, 2026

Run Gemma 4 26B on GTX 1080 with llama.cpp

A developer how-to demonstrating how to run Google’s Gemma 4 26B‑A4B Mixture‑of‑Experts model locally on an 8 GiB NVIDIA GeForce GTX 1080 using an enhanced llama.cpp fork (AtomicBot-ai/atomic-llama-cpp-turboquant). The guide details system setup (driver pinning, CUDA nvcc, gcc-14 workaround, glibc patch), building the fork with CUDA, downloading the main GGUF and MTP assistant head, and tuning offload parameters. Key optimisations include keeping most MoE expert weights in host RAM (streamed over PCIe), using RotorQuant/TurboQuant KV cache to enable 128k context, and forcing the assistant embedding table onto the GPU with --override-tensor-draft to enable effective MTP speculative decoding. The author reports ~24.5 tokens/sec at 128k context and describes the memory/PCIe tradeoffs and the final recommended command-line configuration.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.