In plain words: Safety checks judge requests before seeing how answers will be used, so attackers can copy a benign chat history and pass. Attackers are still guaranteed some help: useful answers, safety, and open access cannot coexist, and only hard-to-copy proof of use lowers it.
Abstract
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.
Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu
arXiv:2607.27951 · cs.CR, cs.AI · submitted Jul 30, 2026
abstract · pdf · html