about
Fast and Optimal Mapping for Accelerator Modeling and Evaluation (arxiv.org)
2 points by PaulHoule 222 days ago | hide | past | pdf | discuss on HN

In plain words: A tool that schedules how a chip runs a neural network by noting where data sits, letting it throw out useless schedules and check the rest. Unlike guess-and-check searches, it guarantees the best schedule and finds it in 17 seconds instead of 5 hours.

Abstract · The Turbo-Charged Mapper: Fast and Optimal Mapping for Energy-efficient and Low-latency Accelerator Design

The energy and latency of an accelerator running a deep neural network (DNN) depend on how the computation and data movement are scheduled in the accelerator (i.e., mapping), and picking an optimal mapping is essential to achieve high-performance accelerators. However, it is challenging to find mappings that maximize accelerator performance. The space of mappings is large, and prior works cannot guarantee finding optimal mappings because they use heuristics or metaheuristics to narrow the search space. To address this challenge, we propose the Turbo-Charged Mapper (TCM), a fast mapper that finds optimal mappings. The key to our approach is that we define a new mapping concept called dataplacement, which, like the prior concept of dataflow, allows for clear analysis and comparison of mappings. Through it, we identify opportunities to prune redundant and suboptimal mappings, reducing search space by up to 32 orders of magnitude ($10^{37}\rightarrow10^5$). TCM leverages these insights to perform full mapspace searches, making it the first mapper that can find optimal mappings in feasible runtime. Compared to prior mappers, TCM improves accelerator energy-delay-product by $1.2-6.5\times$ while simultaneously reducing mapping search time by $1000\times$ (5 hours $\rightarrow$ 17 seconds).

Michael Gilbert, Tanner Andrulis, Vivienne Sze, Joel S. Emer
arXiv:2602.15172 · cs.AR · submitted Feb 16, 2026 · updated May 2, 2026
abstract · pdf · html

add comment on HN