about
Iarpa Trojans in Artificial Intelligence (arxiv.org)
2 points by hlynurd 153 days ago | hide | past | pdf | discuss on HN

In plain words: A multi-year program studied hidden backdoors in AI models that make them fail or let an attacker take over. It built detectors that inspect a model's inner numbers or rebuild the cue; tests showed how well they work and that some backdoors arise naturally.

Abstract · Trojans in Artificial Intelligence (TrojAI) Final Report

The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, hidden backdoors intentionally embedded within an AI model that can cause a system to fail in unexpected ways, or allow a malicious actor to hijack the AI model at will. This multi-year initiative helped to map out the complex nature of the threat, pioneered foundational detection methods, and identified unsolved challenges that require ongoing attention by the burgeoning AI security field. This report synthesizes the program's key findings, including methodologies for detection through weight analysis and trigger inversion, as well as approaches for mitigating Trojan risks in deployed models. Comprehensive test and evaluation results highlight detector performance, sensitivity, and the prevalence of "natural" Trojans. The report concludes with lessons learned and recommendations for advancing AI security research.

Kristopher W. Reese, Taylor Kulp-McDowall, Michael Majurski, Tim Blattner, Derek Juba, Peter Bajcsy, Antonio Cardone, Philippe Dessauw, Alden Dima, Anthony J. Kearsley, Melinda Kleczynski, Joel Vasanth, et al.
arXiv:2602.07152 · cs.CR, cs.AI, cs.LG · submitted Feb 6, 2026 · updated Feb 27, 2026
abstract · pdf

add comment on HN