In plain words: A new open-source compiler turns Triton kernels and PyTorch models straight into runnable code for Qualcomm's Hexagon NPU. It fuses work into one big kernel that keeps data in the chip's fast local memory, cutting the traffic jams that slow down library-based approaches.
Abstract · Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
In this paper, we present Hexagon-MLIR,an open-source compilation stack that targets Qualcomm Hexagon Neural Processing Unit (NPU) and provides unified support for lowering Triton kernels and PyTorch models . Built using the MLIR framework, our compiler applies a structured sequence of passes to exploit NPU architectural features to accelerate AI workloads. It enables faster deployment of new Triton kernels (hand-written or subgraphs from PyTorch 2.0), for our target by providing automated compilation from kernel to binary. By ingesting Triton kernels, we generate mega-kernels that maximize data locality in the NPU's Tightly Coupled Memory (TCM), reducing the bandwidth bottlenecks inherent in library-based approaches. This initiative complements our commercial toolchains by providing developers with an open-source MLIR-based compilation stack that gives them a path to advance AI compilation capabilities through a more flexible approach. Hexagon-MLIR is a work-in-progress, and we are continuing to add many more optimizations and capabilities in this effort.
Mohammed Javed Absar, Muthu Baskaran, Abhikrant Sharma, Abhilash Bhandari, Ankit Aggarwal, Arun Rangasamy, Dibyendu Das, Fateme Hosseini, Franck Slama, Iulian Brumar, Jyotsna Verma, Krishnaprasad Bindumadhavan, et al.
arXiv:2602.19762 · cs.PL, cs.AI · submitted Feb 23, 2026
abstract · pdf · html
The GitHub for hexagon-mlir, https://github.com/qualcomm/hexagon-mlir
Lovely to think such a powerful energy efficient number cruncher might be available for broader use. It feels like this has been nigh impossible to make good use of; maybe incorrect but my impression is Qualcomm has never wanted the hexagon to be in the spotlight, has positioned it as something your BSP maybe used for you. Hexagon-mlir feels like a very necessary renormalization, is a rallying call to get folks interested & using this thing that for decades (the microarchitecture is from 2006, happy birthday) has been there but invisible.