about
SRM: Detecting slow-burn risk in AI-agent sessions before execution (arxiv.org)
1 point by ilion_identity 137 days ago | hide | past | pdf | discuss on HN

In plain words: A small memory module tracks an agent's whole session, not just each action alone, so attacks spread across many harmless-looking steps get caught. On 80 sessions it flagged every attack and dropped false alarms from 5% to 0, adding under 250 microseconds per step.

Abstract · Session Risk Memory (SRM): Temporal Authorization for Deterministic Pre-Execution Safety Gates

Deterministic pre-execution safety gates evaluate whether individual agent actions are compatible with their assigned roles. While effective at per-action authorization, these systems are structurally blind to distributed attacks that decompose harmful intent across multiple individually-compliant steps. This paper introduces Session Risk Memory (SRM), a lightweight deterministic module that extends stateless execution gates with trajectory-level authorization. SRM maintains a compact semantic centroid representing the evolving behavioral profile of an agent session and accumulates a risk signal through exponential moving average over baseline-subtracted gate outputs. It operates on the same semantic vector representation as the underlying gate, requiring no additional model components, training, or probabilistic inference. We evaluate SRM on a multi-turn benchmark of 80 sessions containing slow-burn exfiltration, gradual privilege escalation, and compliance drift scenarios. Results show that ILION+SRM achieves F1 = 1.0000 with 0% false positive rate, compared to stateless ILION at F1 = 0.9756 with 5% FPR, while maintaining 100% detection rate for both systems. Critically, SRM eliminates all false positives with a per-turn overhead under 250 microseconds. The framework introduces a conceptual distinction between spatial authorization consistency (evaluated per action) and temporal authorization consistency (evaluated over trajectory), providing a principled basis for session-level safety in agentic systems.

Florin Adrian Chitan
arXiv:2603.22350 · cs.AI, cs.CR · submitted Mar 22, 2026
abstract · pdf · 12 pages, 3 figures. Companion paper to arXiv:2603.13247. Benchmark dataset and artifacts available on Zenodo: 10.5281/zenodo.15410944

add comment on HN