about
Leace: Perfect linear concept erasure in closed form (arxiv.org)
5 points by anigbrowl on Jun 7, 2023 | hide | past | pdf | 1 comment on HN

In plain words: LEACE gives one formula that edits an embedding so no straight-line classifier can spot the concept, while changing the vector as little as possible. Applied to every layer of a language model, it reduced gender bias in BERT embeddings.

Abstract · LEACE: Perfect linear concept erasure in closed form

Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called "concept scrubbing," which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Code is available at https://github.com/EleutherAI/concept-erasure.

Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, Stella Biderman
arXiv:2306.03819 · cs.LG, cs.CL, cs.CY · submitted Jun 6, 2023 · updated Apr 3, 2025
abstract · pdf · html

add comment on HN

I only vaguely understand what this is or what it's for, but I love it because it's probably the most SCP sounding paper title I've seen IRL