about
BiT: Robustly Binarized Multi-Distilled Transformer (arxiv.org)
2 points by jasondavies on Sep 6, 2023 | hide | past | pdf | discuss on HN

In plain words: It squeezes a language model by storing every weight and activation as just +1 or −1, then trains it step by step to copy more precise versions of itself. The shrunken model scored within 5.9% of the usual full-precision one on a language-understanding test.

Abstract · BiT: Robustly Binarized Multi-distilled Transformer

Modern pre-trained transformers have rapidly advanced the state-of-the-art in machine learning, but have also grown in parameters and computational complexity, making them increasingly difficult to deploy in resource-constrained environments. Binarization of the weights and activations of the network can significantly alleviate these issues, however, is technically challenging from an optimization perspective. In this work, we identify a series of improvements that enables binary transformers at a much higher accuracy than what was possible previously. These include a two-set binarization scheme, a novel elastic binary activation function with learned parameters, and a method to quantize a network to its limit by successively distilling higher precision models into lower precision students. These approaches allow for the first time, fully binarized transformer models that are at a practical level of accuracy, approaching a full-precision BERT baseline on the GLUE language understanding benchmark within as little as 5.9%. Code and models are available at: https://github.com/facebookresearch/bit.

Zechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao, Scott Yih, Meng Li, Raghuraman Krishnamoorthi, Yashar Mehdad
arXiv:2205.13016 · cs.LG, cs.CL · submitted May 25, 2022 · updated Oct 2, 2022
abstract · pdf · html · NeurIPS 2022

add comment on HN