In plain words: A new design language lets an AI describe and speed-test inference system designs without working code, so it can invent architectures instead of tuning existing software. An agent found designs improving speed and responsiveness by 6.23-50.1% over the best setup for DeepSeek V4 Pro.
Abstract
AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based. Profiling-bound feedback confines the search to the capabilities and performance of an existing software stack, preventing a fundamentally better architecture of LLM inference systems from being identified. To enable the AI-driven LLM inference system architecting loop, we argue that a general workload representation, a verifiable mutation space, and an implementation-independent evaluator are required. We present the RoofLang domain-specific language (DSL) that provides these features. In our evaluation, RoofLang reveals that DeepSeek V4-series models could achieve 3.5-39.5$\times$ higher peak decode throughput than other representative models. This gap is disproportionate to their total parameter counts and arises largely from compact KV-cache designs that support larger batches and reduce memory traffic. A persistent optimizer agent further discovered several new architectures that improved both throughput and interactivity of DeepSeek V4 Pro on NVIDIA B300 by 6.23-50.1%.
Ziyue Yang, Yuting Jiang, Lei Qu, Peng Cheng
arXiv:2609.12551 · cs.DC, cs.AI · submitted Sep 11, 2026 · updated Sep 20, 2026
abstract · pdf · html · v2: updated author affiliations and added a missing statement in Section 5's "Modeling fidelity" paragraph