In plain words: It finds and edits small groups of connections inside trained neural networks, so researchers can test which parts handle which jobs. The free toolkit plugs into widely used model code, letting people carve up networks without building such tools themselves.
Abstract
Despite recent advances in the field of explainability, much remains unknown about the algorithms that neural networks learn to represent. Recent work has attempted to understand trained models by decomposing them into functional circuits (Csordás et al., 2020; Lepori et al., 2023). To advance this research, we developed NeuroSurgeon, a python library that can be used to discover and manipulate subnetworks within models in the Huggingface Transformers library (Wolf et al., 2019). NeuroSurgeon is freely available at https://github.com/mlepori1/NeuroSurgeon.
Michael A. Lepori, Ellie Pavlick, Thomas Serre
arXiv:2309.00244 · cs.LG, cs.CL · submitted Sep 1, 2023
abstract · pdf · html