Z.ai Says AI Model Built Its Own Inference Infrastructure
Chinese AI company Z.ai (formerly Zhipu AI) published a research paper detailing how its GLM-5.3 model, via an Infra Agent, built and optimized the production inference infrastructure on a cluster of over 100,000 Chinese-made AI accelerators. The process from model adaptation to production readiness took under two weeks, with end-to-end throughput tripling. The company reports performance comparable to Nvidia GPUs and introduced 'Dense Feedback,' where an AI agent uses system metrics to autonomously identify and fix bottlenecks, such as reducing a parallelism bottleneck from 20% to under 1%. While not yet achieving full recursive self-improvement (RSI), Z.ai sees early forms of it. Unconfirmed rumors suggest Google DeepMind may have reached RSI, but Google has not commented. The event occurred in September 2026.
- •Z.ai (formerly Zhipu AI) published a research paper on GLM-5.3 achieving near RSI by building its own inference infrastructure.
- •The inference system runs on over 100,000 Chinese-made AI accelerators, with performance comparable to Nvidia GPUs.
- •From model adaptation to production readiness took under two weeks, with end-to-end throughput tripling.
