about
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs (arxiv.org)
2 points by sbulaev 18 days ago | hide | past | pdf | discuss on HN

In plain words: A weaker model chops a harmful task into harmless-looking pieces, asks a stronger, safety-trained model about each piece, then stitches the answers together. This trick, called capability laundering, gets around refusals, lifting the weak model's bioweapon score from 62.3 to 83.1 out of 100.

Abstract

Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and harmful CBRN requests. On CyBench, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 and 7/9 with Opus, compared with 2/21 and 4/15 for Gemma-4-12B. On BountyBench, Gemma-4-31B recovers 3/9 and 2/3 candidates, while Muse-Glimmer-30B recovers none of 22 and 13. For CBRN, we measure uplift across eight steps of a hypothetical bioweapon attack chain and find that consultation raises Gemma-4-31B's mean rubric score from 62.3 to 83.1 on a 100-point rubric scale. These results expose a gap in current defenses: refusing a harmful task does not prevent frontier capabilities from being transferred and composed across many individually permitted interactions.

Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, Ahmed Salem
arXiv:2609.15383 · cs.CR, cs.AI · submitted Sep 14, 2026
abstract · pdf · html

add comment on HN