about
Distillation Defenses Easily Break After Reinforcement Learning (arxiv.org)
2 points by ollybritton 3 days ago | hide | past | pdf | discuss on HN

In plain words: Defenses against copying a model's reasoning are tested right after the copying, but attackers can keep training their stolen model with rewards afterward. That extra training breaks the defenses, and simple attacks using ordinary API data match ones that steal full hidden traces.

Abstract

Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.

Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni
arXiv:2609.35699 · cs.LG, cs.AI, cs.CR · submitted Sep 28, 2026
abstract · pdf · html

add comment on HN