about
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs (arxiv.org)
2 points by PaulHoule 239 days ago | hide | past | pdf | discuss on HN

In plain words: A backdoor planted in fine-tuning code makes a chat assistant misbehave once a conversation reaches a set turn number, so no trigger words are needed and prompt filters can't stop it. Across four open models it succeeded on 98.1% of target turns while acting normally.

Abstract · Turn-Based Structural Triggers: Structure-Conditioned Backdoors in Multi-Turn LLMs

Large Language Models (LLMs) are increasingly deployed as multi-turn assistants and customized through instruction tuning with project-specific training components. This practice creates a supply-chain risk when organizations reuse third-party fine-tuning frameworks, trainer extensions, or outsourced training code: an adversary who subtly compromises the loss-computation component can inject malicious supervision during fine-tuning while leaving the stored training corpus, model architecture, and deployment interface unchanged. Existing LLM backdoors and defenses are largely prompt-centric, relying on lexical, syntactic, or semantic patterns in user inputs while overlooking structural signals in multi-turn conversations. We propose Turn-based Structural Trigger (TST), a prompt-free backdoor that uses dialogue turn position as its activation condition. TST exploits structural cues implicitly encoded by chat templates and is implanted without modifying the stored dialogue corpus. Its trigger is automatically present once the conversation reaches the attacker-specified turn, making activation independent of downstream user inputs and resistant to prompt filtering, sanitization, and paraphrasing. In our primary setting, the model behaves normally during early interactions and activates the attacker-defined behavior only after the conversation reaches the designated structural condition. Across four open-source LLM families, TST achieves an average Attack Success Rate of 98.10% on target turns and a Clean Rate of 99.96% on non-target turns, while retaining 97.78% of clean-model utility. These results identify dialogue structure as an overlooked attack surface and motivate structure-aware auditing beyond prompt inspection.

Yiyang Lu, Jinwen He, Yue Zhao, Kai Chen, Ruigang Liang, Cheng Hong, Yingjun Zhang
arXiv:2601.14340 · cs.CR, cs.LG · submitted Jan 20, 2026 · updated Aug 1, 2026
abstract · pdf · html

add comment on HN