Observed Signal · Aug 10, 2026 · Technical Release · Source: SemiAnalysis · Impact: 3/5 · Sentiment: Positive
TileRT Enables Ultra-High Interactivity on NVIDIA GPUs
SemiAnalysis reports on TileRT, a persistent-engine approach that compiles an entire decode graph into a single resident kernel on NVIDIA GPUs to reduce per-token latency. In InferenceX benchmarks TileRT reached up to 500 tokens/s/user on a single B200 decode server (GLM5 FP8 744B) and delivered large gains versus traditional GPU inference engines (e.g., ~3× vs GB300 NVL72 in some tests). TileRT is designed to handle latency-sensitive decode while remaining interoperable with throughput-optimized prefill engines such as vLLM. TileRT is already deployed in production at Xiaomi and Z.ai, but currently supports a small model catalog (GLM-5/5.1, DeepSeek-V3.2) and primarily serves batch size 1 decode workloads, reflecting trade-offs between per-user interactivity and aggregate throughput.
Benchmarks and a software approach that materially improve GPU per-user latency (and are already in production) can shift operational choices between fungible GPU fleets and purpose-built dataflow ASICs; this affects LLM serving economics and capacity planning for AI services.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- TileRT statically compiles the entire decode graph into a single persistent kernel on NVIDIA GPUs to maximize overlap between computation, memory, and communication.
- InferenceX benchmark: TileRT reached up to 500 tokens/s/user on a single B200 decode server (GLM5 FP8 744B), approximately 3× faster than GB300 NVL72 running traditional inference engines.
- On an eight-GPU B200 node TileRT reached 340 tokens/s/user (8k/1k scenario) and at 1k/1k TileRT FP8 reached 494.2 tokens/s/user in the reported dataset.
- TileRT is deployed in production at Xiaomi (MiMo V2.5 Pro UltraSpeed) and at Z.ai for GLM-5.1 HighSpeed.
- TileRT currently supports a limited model catalog (GLM-5/5.1 and DeepSeek-V3.2) and, as of publication, each decode node serves one in-flight request (batch size 1), trading aggregate throughput for per-user latency.
Connected Companies & Entities
18 Entities mapped“Frontier AI labs such as OpenAI are therefore evaluating purpose-built inference systems, including Cerebras and NVIDIA Groq LPUs that prior...”
“We thank the TileRT maintainers for collaborating on TileRT InferenceX benchmarks and also in general thankful to the vLLM community for the...”
“Frontier AI labs such as OpenAI are therefore evaluating purpose-built inference systems, including Cerebras and NVIDIA Groq LPUs that prior...”
“Frontier AI labs such as OpenAI are therefore evaluating purpose-built inference systems, including Cerebras and NVIDIA Groq LPUs that prior...”
“Frontier AI labs such as OpenAI are therefore evaluating purpose-built inference systems, including Cerebras and NVIDIA Groq LPUs that prior...”
“TileRT does not replace vLLM, vLLM remains the high-throughput prefill engine and the surrounding serving layer......”
“The TileRT decode engine is already being deployed in production at Xiaomi for MiMo V2.5 Pro UltraSpeed......”
“The TileRT decode engine is already being deployed in production at Xiaomi for MiMo V2.5 Pro UltraSpeed and ZAI with GLM 5.1 HighSpeed....”
“The TileRT decode engine is already being deployed in production at Xiaomi for MiMo V2.5 Pro UltraSpeed and ZAI with GLM 5.1 HighSpeed....”
“Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft A...”
“Furthermore, it has the support of the ML community ... and the support of major labs like OpenAI, MiniMax, ZAI, Qwen, Moonshot Kimi, etc....”
“Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft A...”
“Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft A...”
“We will have Google TPUv7 results soon, and AMD has committed to MI455X UALoE72 this year too....”
“Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft A...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
TileLang for High-Performance GPU Kernels
This technical tutorial introduces TileLang, a Python-first DSL for writing high-performance GPU kernels that balances the high-level convenience of Triton with the low-level control of CUTLASS/CuTe. TileLang exposes tiles as first-class objects, requires explicit buffer placement (shared memory, registers), and relies on a layout-inference pass to derive thread mappings and memory layouts. The article walks through a GEMM example, an MLA (Multi‑Head Latent Attention) decode kernel (DeepSeek) where TileLang's layout inference enables a compact ~80-line implementation matching FlashMLA H100 fp16 performance, and a production RMSNorm+SiLU drop-in used at AtlasCloud that expands supported channel widths and improved latency. The post details primitives (T.alloc_shared, T.gemm, T.Pipelined, etc.), backend targets, and one-line optimization knobs like swizzling and warp policies.
cuTile Rust: Safe Rust GPU Kernels at Near-cuBLAS Speed
cuTile Rust, a tile-based DSL and crate introduced by NVIDIA researchers in the paper “Fearless Concurrency on the GPU” (arXiv:2606.15991), applies Rust’s ownership and borrow-checker model across the host-to-GPU launch boundary by partitioning mutable outputs into provably disjoint tiles and passing exclusive &mut views to tile kernels. The approach compiles to CUDA Tile IR and then into GPU cubins, requiring sm_80+ GPUs, CUDA 13.3, Rust 1.89+, and Linux. Authors report throughput reaching about 96% of cuBLAS on GEMM (on a B200) and end-to-end Grout inference results (171 tok/s for Qwen3-4B on an RTX 5090), though independent reproduction varies by hardware and workload. The crate and toolchain are early-stage, CUDA/Linux-only, and API/macros may change between releases.
Google TPUv7 Ironwood Delivers Up to 50% Better Performance per Dollar
SemiAnalysis, an independent analyst firm, published third-party inference benchmarks for Google's TPUv7 Ironwood, comparing it against NVIDIA's B200 and B300 GPUs. The results show Ironwood delivering up to 50% better performance per dollar in apples-to-apples comparisons, with a 19-34% lower cost per token at typical interactivity levels. Key factors include Google's co-designed hardware and software stack, the emerging TorchTPU external stack providing native PyTorch support, and optimizations across kernels and serving engines (vLLM, SGLang). Google is externalizing its TPU infrastructure, selling chips outright and on Google Cloud, with Anthropic as a major customer. The TorchTPU stack is in private beta, open-sourcing around mid-October, and future work includes speculative decoding, disaggregated serving, and KV-cache offloading, positioning TPUs as a strong competitor in the AI inference market.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
