about
Multi-Turn Intent Detection for LLM and Agent Security (ArXiv) (arxiv.org)
1 point by sharathr 226 days ago | hide | past | pdf | 1 comment on HN

In plain words: A safety monitor that remembers the whole chat, spotting threats that build slowly across messages instead of judging each one alone. It scored 0.84 on the standard detection score versus 0.67 for the best current filters, adding under 20 milliseconds per reply.

Abstract · DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in LLMs

While Large Language Model (LLM) capabilities have scaled, safety guardrails remain largely stateless, treating multi-turn dialogues as a series of disconnected events. This lack of temporal awareness facilitates a "Safety Gap" where adversarial tactics, like Crescendo and ActorAttack, slowly bleed malicious intent across turn boundaries to bypass stateless filters. We introduce DeepContext, a stateful monitoring framework designed to map the temporal trajectory of user intent. DeepContext discards the isolated evaluation model in favor of a Recurrent Neural Network (RNN) architecture that ingests a sequence of fine-tuned turn-level embeddings. By propagating a hidden state across the conversation, DeepContext captures the incremental accumulation of risk that stateless models overlook. Our evaluation demonstrates that DeepContext significantly outperforms existing baselines in multi-turn jailbreak detection, achieving a state-of-the-art F1 score of 0.84, which represents a substantial improvement over both hyperscaler cloud-provider guardrails and leading open-weight models such as Llama-Prompt-Guard-2 (0.67) and Granite-Guardian (0.67). Furthermore, DeepContext maintains a sub-20ms inference overhead on a T4 GPU, ensuring viability for real-time applications. These results suggest that modeling the sequential evolution of intent is a more effective and computationally efficient alternative to deploying massive, stateless models.

Justin Albrethsen, Yash Datta, Kunal Kumar, Sharath Rajasekar
arXiv:2602.16935 · cs.AI, cs.ET, cs.LG · submitted Feb 18, 2026
abstract · pdf · html · 18 Pages, 7 Tables, 1 Figure

add comment on HN

Hi HN — I’m one of the authors.

We’ve been working on security for multi-turn agent loops and noticed most detection approaches operate on isolated prompts. This paper introduces a framework for modeling intent trajectories across sequences in real time (<20ms), enabling enforcement before harmful actions occur.

Happy to answer technical questions or discuss assumptions in the paper.