In plain words: A C++ toolkit for GPU math that copies the look of the popular CPU library Armadillo, so old code moves over easily. It merges each chain of operations into one GPU job at compile time, beating PyTorch, TensorFlow, and JAX, sometimes by a lot.
Abstract
We describe the Bandicoot GPU linear algebra toolkit for C++, which prioritises ease of use without compromising efficiency. Bandicoot's API aims for compatibility with the popular Armadillo CPU linear algebra library, enabling easy transition for existing CPU-based codebases. Unlike other GPU-focused toolkits, Bandicoot uses template metaprogramming to generate fused GPU kernels directly at compile-time, yielding efficient kernels that can saturate memory bandwidth. This removes the need for run-time overhead or JIT infrastructure. Empirical results show that Bandicoot outperforms (sometimes by considerable margins) commonly-used linear algebra toolkits including PyTorch, TensorFlow, and JAX.
Ryan R. Curtin, Marcus Edel, Conrad Sanderson
arXiv:2604.22242 · cs.MS · submitted Apr 24, 2026 · updated Jul 3, 2026
abstract · pdf · html