In plain words: A new rule for tracking how outputs change lets networks with only true/false weights and inputs train using logic gates instead of gradient math and real numbers. It matched full-precision accuracy on ImageNet classification while cutting energy use in both training and inference.
Abstract · BOLD: Boolean Logic Deep Learning
Deep learning is computationally intensive, with significant efforts focused on reducing arithmetic complexity, particularly regarding energy consumption dominated by data movement. While existing literature emphasizes inference, training is considerably more resource-intensive. This paper proposes a novel mathematical principle by introducing the notion of Boolean variation such that neurons made of Boolean weights and inputs can be trained -- for the first time -- efficiently in Boolean domain using Boolean logic instead of gradient descent and real arithmetic. We explore its convergence, conduct extensively experimental benchmarking, and provide consistent complexity evaluation by considering chip architecture, memory hierarchy, dataflow, and arithmetic precision. Our approach achieves baseline full-precision accuracy in ImageNet classification and surpasses state-of-the-art results in semantic segmentation, with notable performance in image super-resolution, and natural language understanding with transformer-based models. Moreover, it significantly reduces energy consumption during both training and inference.
Van Minh Nguyen, Cristian Ocampo, Aymen Askri, Louis Leconte, Ba-Hien Tran
arXiv:2405.16339 · stat.ML, cs.LG · submitted May 25, 2024 · updated Jun 6, 2025
abstract · pdf · html · Published at NeurIPS 2024 main conference
But their framework is general and compatible with binary, ternary an floating point precision. So they allow for training mixed precision models, which is probably important for models which require high precision in some aspects to achieve high downstream accuracy. (As an example they cite image segmentation and super resolution, for which they also include benchmarks to demonstrate achieving high accuracy.)
The whole thing could substantially reduce training time, apart from inference time and memory requirements. (Most conventional BNNs only reduce the latter two, but still require training a full precision model first. Or when training from scratch, they still use floating point operations in various ways. Which probably explains why conventional BNNs are currently not used for training.)
The main limitation of their work is the fact that current GPUs aren't optimized for binary operations, only for floating point operations (mainly multiplication). So the authors calculate analytically how much their approach would reduce energy consumption during training, which should closely correspond to actual "compute" requirements on optimized hardware. For their benchmarks they show the calculated energy requirement to be only a fraction of the floating point base line (and even of other conventional BNNs) while achieving similar model accuracy.
It's also interesting to note that this is an academic publication (it was accepted for NeurIPS 2024) but sponsored by Huawei. So the full code is not available, though they provide a lot of implementation details in the paper/appendix. I wonder whether Huawei jumps on the opportunity and develops a machine learning accelerator for binary operations, which would compete with GPUs that are mainly optimized for FLOPs.