about
A Barrier-Free Synchronization Algorithm for Multi-Engine AI Accelerators (arxiv.org)
1 point by matt_d 47 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of halting every engine at each loop turn, the compiler tracks loop counts as the program runs so each task waits only for the exact data it needs. On ML kernels this cut latency 10-45% versus the stop-everything approach.

Abstract

Multi-engine AI accelerators such as AWS Trainium comprise specialized compute engines that execute in parallel, and the compiler must synchronize the data dependencies between them. For straight-line code this is simple: each dependency reduces to waiting for a threshold count of instruction completions, which the compiler computes statically. Loops admit no such static threshold; a simple solution inserts all-engine barriers at iteration boundaries, resetting synchronization state so each loop body can be treated as straight-line, at the cost of parallelism. We present a barrier-free synchronization algorithm that instead enforces each dependency precisely across structured control flow with arbitrarily nested, dynamically bounded loops. The key idea is to compute dynamic thresholds at runtime from tracked loop iteration counts. We implemented it as a compiler backend pass at the AWS Neuron ISA level. On a suite of ML kernels, it reduces latency 10-45% relative to the barrier-based baseline, achieves a 3.3x speedup on a synchronization-bound microbenchmark, and often matches or exceeds hand-tuned manual allocation. Issuing a consumer too early violates its dependency, while issuing too late unnecessarily stalls execution. We formally characterize the minimum synchronization required for correctness and verify in the Lean proof assistant, via bisimulation, that our algorithm meets this criterion.

Chungha Sung, Nikil V. Shyamsunder, Hanliang Zhang, Daniel Kroening, Joonwon Choi
arXiv:2608.13757 · cs.PL, cs.DC · submitted Aug 13, 2026
abstract · pdf · html · To appear in the 2027 IEEE/ACM International Symposium on Code Generation and Optimization (CGO '27). 16 pages, 8 figures, 2 tables

add comment on HN