In plain words: Instead of letting whole GPU compute and communication jobs run side by side, this tool breaks communication into small chunks and weaves them into one GPU program as data arrives. It sped up multi-GPU jobs 1.3 times on average, versus the usual whole-job overlap.
Abstract · Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
Communication has become a first-order bottleneck in large-scale GPU workloads, and existing distributed compilers address it mainly by overlapping whole compute and communication kernels at the stream level. This coarse granularity incurs extra kernel launches, forces device-wide synchronizations at kernel boundaries, and leaves substantial slack when the slowest tile or kernel stretches the communication tail. We present Syncopate, a compiler and runtime that enables automatic fine-grained overlap inside a single fused kernel. Syncopate introduces a communication chunk abstraction that decouples communication granularity from kernel structure and backend mechanisms, allowing chunk-level plans to be ported from existing distributed compilers, written directly by users, or instantiated from reusable templates. Given a local Triton kernel and a chunk schedule, Syncopate performs transformations to align computation with chunk availability. Implemented as a source-to-source compiler on Triton, Syncopate delivers an average end-to-end speedup of 1.3$\times$ and up to 4.7$\times$ on multi-GPU workloads. Our code is open-sourced at https://github.com/tie-pilot-qxw/syncopate.
Xinwei Qiang, Yue Guan, Zhengding Hu, Keren Zhou, Yufei Ding, Adnan Aziz
arXiv:2601.20595 · cs.DC · submitted Jan 28, 2026 · updated Jul 1, 2026
abstract · pdf · html · Accepted at the 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2026). Camera-ready version with artifact appendix