about
PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch (arxiv.org)
84 points by mfiguiere on Apr 24, 2025 | hide | past | pdf | 8 comments on HN

In plain words: It rewrites programs so their GPU work can be recorded once and replayed as a single CUDA Graph, skipping the CPU cost of launching each tiny kernel. Across many ML workloads, this more than doubled the speedup from graph replay compared with PyTorch's built-in support.

Abstract

Machine learning (ML) workloads launch hundreds to thousands of short-running GPU kernels per iteration. With GPU compute throughput growing rapidly, CPU-side launch latency of kernels is emerging as a bottleneck. CUDA Graphs promise to address this by replaying a set of kernels with a single dispatch of the graph, removing per-kernel launch costs. However, CUDA Graphs remain surprisingly difficult to deploy correctly and efficiently. We present PyGraph - a compiler framework to maximize the coverage and benefits of CUDA Graphs for ML workloads. It introduces three novel optimizations: it applies automatic code transformations to make ML applications amenable to CUDA Graphs; it eliminates the parameter copy overheads for kernels executing in CUDA Graphs, and it selectively deploys CUDA Graphs guided by a cost-benefit analysis. For 25 ML workloads from TorchBench, HuggingFace, and TIMM, PyGraph more than doubles the benefit from deploying CUDA Graph compared to the most popular and widely used ML compiler, PyTorch2. PyGraph is built atop PyTorch2's compilation framework and requires no programmer intervention.

Abhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava Basu
arXiv:2503.19779 · cs.LG · submitted Mar 25, 2025 · updated Dec 22, 2025
abstract · pdf · html

add comment on HN

Python can be used for many types of graphs. This package is for CUDA Graphs, so wouldn't "PyCudaGraph" be a better name?
The lack of a readily available, installable package (pip install pygraph - has no relation to this paper as far as i can tell) makes it difficult to fully assess the reproducibility and practical applicability of the work.
why request code.. when all of pytorch2 is open, and this is built on top of it with some enhancements, why not put this out in the open
I think that might be just a feature of the catalyzex platform for papers with no linked code yet that might internally add a +1 to code requested on their db and thats it

some times papers come out a few weeks before code when its bleeding edge

This is neat, although it would be nice to see it merged into PyTorch instead of just a paper :) The key seems to be (beyond "obvious" optimizations like not running graphs that are measured to be slower) is that graphs "bake-in" parameters and if those change then the graph needs to be thrown away. The solution is indirecting more, so that what gets captured is a pointer that can remain constant, while the data behind it is changed. This also saves the need to copy in and out of a graph-captured buffer because you can just swap out the pointer instead. Of course there is overhead to this approach (I don't think the authors actually explore this much) in that you throw away information (divisibility, for example) that would allow for constructing better kernels, but often this is still worth it. (Or you could pass this through too.)

Something worth exploring later would be getting better support for the rest of CUDA graphs into PyTorch, like conditional nodes.

Nice to see work by IISc show up on HN.

Uday Bondhugula, the lead developer of Pluto framework for polyhedral comp. is also at IISc, whose group has spun out a startup,

https://www.polymagelabs.com/

Nice to see IISc support cool stuff like this (incl. their ArtPark initiative.)

I don't see any source code.