In plain words: A survey of how to split work between analog memory tiles and digital units turns scattered approaches into one four-stage plan, tested on GPT-2. It finds 4 of 49 projections dominate the errors, the first block's attention output ten times more sensitive than the rest.
Abstract · Heterogeneous Mapping for Analog In-Memory Computing Accelerators: A Unified Workflow
Analog In-Memory Computing (AIMC) accelerators execute matrix-vector multiplications directly within memory arrays, reducing data movement and improving DNN inference efficiency. Their limited effective precision motivates heterogeneous architectures that combine analog compute tiles with digital processing units. This letter classifies existing methods for partitioning DNN workloads across these resources by mapping granularity, optimization strategy, and model support, and distills them into a unified four-stage workflow. To demonstrate the workflow on a model class not yet addressed by existing methods, we apply its first two stages to GPT-2, producing the first AIMC-specific precision sensitivity profile for a decoder-only transformer. Sensitivity is dominated by 4 of 49 projections, with the first decoder block's attention output dominating by an order of magnitude. This suggests that projection-level mapping and selective digital execution of early-block and output-facing projections are important for reliable decoder-transformer deployment on AIMC hardware.
Corey Lammie
arXiv:2606.02672 · cs.AR, cs.ET · submitted Jun 1, 2026
abstract · pdf · html · Accepted by IEEE Computer Architecture Letters