about
Can we bootstrap AI Safety despite being unable to even define it? (arxiv.org)
2 points by cryptohell 325 days ago | hide | past | pdf | 2 comments on HN

In plain words: Instead of blending several AI models' outputs, this draws samples only where most agree and refuses otherwise, so one unsafe model can't poison it. Its risk stays close to the average of the safest few models, the best tradeoff with how often it abstains.

Abstract · Consensus Sampling for Safer Generative AI

Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions to boost safety. We present consensus sampling, a black-box algorithm that, given k distributions, has risk competitive with the average risk of the safest $s$ while abstaining when there is insufficient agreement. This yields an architecture-agnostic approach to generative-model safety when the distributions are induced by models that can sample and evaluate output probabilities. We formalize the guarantee through R-robustness, which also bounds information leakage and adversarial influence. Inspired by robust statistics and the provable copyright protection algorithm of Vyas et al (2023), we show that while a standard mixture is vulnerable to one unsafe constituent, a pointwise-median construction provides robust intuition, and our efficient sampler is Pareto-optimal for the tradeoff between worst-case risk and abstention. Experiments on synthetic distributions and image generation illustrate the general mechanism and its motivating safety application. The method requires overlap among safe distributions, but it provides a model-agnostic way to inherit guarantees from an unknown reliable subset.

Adam Tauman Kalai, Yael Tauman Kalai, Or Zamir
arXiv:2511.09493 · cs.AI, cs.LG · submitted Nov 12, 2025 · updated May 9, 2026
abstract · pdf · html

add comment on HN

AI output is modeled on human behavior. Are humans safe?
Given several models, assuming only that some unknown subset is "safe", can we construct a single model as safe as that subset? This reduces obtaining a trustworthy model to a plausibly easier task.