about
Don't Lie to Me: Avoiding Malicious Explanations with Stealth (arxiv.org)
1 point by headalgorithm on Jan 26, 2023 | hide | past | pdf | 2 comments on HN

In plain words: STEALTH groups similar data into clusters, then asks the AI model just one label question per cluster so a dishonest model can't tell when it is being tested. This dodges lying and unfair answers that appear when a model is queried many times.

Abstract · Don't Lie to Me: Avoiding Malicious Explanations with STEALTH

STEALTH is a method for using some AI-generated model, without suffering from malicious attacks (i.e. lying) or associated unfairness issues. After recursively bi-clustering the data, STEALTH system asks the AI model a limited number of queries about class labels. STEALTH asks so few queries (1 per data cluster) that malicious algorithms (a) cannot detect its operation, nor (b) know when to lie.

Lauren Alvarez, Tim Menzies
arXiv:2301.10407 · cs.SE, cs.AI, cs.CR · submitted Jan 25, 2023
abstract · pdf · html · 6 pages, 6 Tables, 3 figures

add comment on HN

mk kjkj
kkmk