about
Reducing Pipeline Bubbles with Adaptive Parallelism on Heterogeneous Models (arxiv.org)
2 points by PaulHoule 355 days ago | hide | past | pdf | 1 comment on HN

In plain words: When an AI model is split across GPUs, some chips sit idle waiting. This system jointly decides how to cut the model, where to put the pieces, and the order to run them, speeding training 1.09 to 1.49 times over the best prior setups.

Abstract · OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling

Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerbates pipeline bubbles, thereby reducing training efficiency. Prior approaches typically optimize a single phase of the pipeline schedule (i.e., partitioning, placement, or scheduling), leaving substantial pipeline bubbles. While promising, co-optimization poses three key challenges: (1) complex performance modeling, (2) a combinatorial search space, and (3) irregular execution orders. To address these challenges, we propose OctoPipe, a pipeline parallelism system to jointly optimize partitioning, placement, and scheduling. First, we build a graph-based pipeline simulator to model heterogeneous pipeline execution for co-optimization. Second, on top of the simulator, we develop an iterative bubble-aware tuner to efficiently explore the combinatorial search space. Third, we implement a unified pipeline executor that dynamically orchestrates computation and communication to support irregular execution orders without deadlocks while maximizing communication-computation overlap. Experiments show that OctoPipe achieves 1.09--1.49$\times$ throughput improvement over the state-of-the-art pipeline parallelism approaches across various heterogeneous model configurations and GPU cluster scales.

Jihu Guo, Tenghui Ma, Wei Gao, Peng Sun, Xun Chen, Jiaxing Li, Zhisheng Ye, Yuyang Jin, Dahua Lin
arXiv:2509.23722 · cs.DC, cs.AI · submitted Sep 28, 2025 · updated Sep 2, 2026
abstract · pdf · html · 14 pages, 13 Figures; Accepted by SC'26;

add comment on HN

Should we ever know what kind of person flagged this

https://news.ycombinator.com/item?id=45580780

:)

It seems to be a take on ROC and could be phrased more humorously. Replacing either acceptance/denial with "filter" makes it not at all insightful,so it is at least a step up from tautology.

In the acceptance case, author presumes(?) that the designer looks only at false positives, etc. there's deep stuff here if people want to dig (postselection fallacies happen in the classical world too!) but poster might have been high on something.

I don't see a world* in which allowing the irreversible flagging of this is positive sum :)

*I only see one in which HN is training some model on the commentariat but not on the "fast-grants" (that's my label for it rn) based policy optimization itself :)

To the posted paper. top schools asking interesting questions, great, if I were to try to make a useful prediction it is that this is going into a second tier conf at best, this seems to be "owned" by junior members of the pjlab team (disregard the corresponding author), is a clear intern project, just a tier below "fast-grants" eligibility.

Chinese schools have weird policy optimization tricks :)