about
Shaping capabilities with token-level data filtering (arxiv.org)
2 points by brandonb 246 days ago | hide | past | pdf | 1 comment on HN

In plain words: Instead of stripping unwanted skills from a finished model, which can be undone, they delete the relevant words from the training text itself. Cutting individual words rather than whole documents hurt normal skills less, and in the biggest models made medical knowledge 7000 times harder to learn.

Abstract

Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple intervention of filtering pretraining data is highly effective, robust, and inexpensive at scale. Inspired by work on data attribution, we show that filtering tokens is more effective than filtering documents, achieving the same hit to undesired capabilities at a lower cost to benign ones. Training models spanning two orders of magnitude, we then demonstrate that filtering gets more effective with scale: for our largest models, token filtering leads to a 7000x compute slowdown on the forget domain. We also show that models trained with token filtering can still be aligned on the forget domain. Along the way, we introduce a methodology for labeling tokens with sparse autoencoders and distilling cheap, high-quality classifiers. We also demonstrate that filtering can be robust to noisy labels with sufficient pretraining compute.

Neil Rathi, Alec Radford
arXiv:2601.21571 · cs.LG, cs.AI, cs.CL · submitted Jan 29, 2026 · updated Jan 30, 2026
abstract · pdf · html · update figure 2

add comment on HN

This is the first new paper from Alec Radford since leaving OpenAI. Token-level data filtering is kind of a simple idea, but so are many effective ideas in LLMs.

One advantage is that this type of safety guardrail can't be undone by an adversary in post-training, so it's a good fit for open source models.

The experiments are all done in preventing models from acquiring medical capabilities, while preserving related capabilities like e.g., biology.