about
Evaluating Gemini Models for Dangerous Capabilities (arxiv.org)
3 points by fishfish on Jun 3, 2024 | hide | past | pdf | discuss on HN

In plain words: Tests were created to check whether an AI can persuade or deceive people, hack computer systems, copy and spread itself, and reason about itself. Tried on a leading AI model, they found no strong dangerous abilities but some early warning signs.

Abstract · Evaluating Frontier Models for Dangerous Capabilities

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, et al.
arXiv:2403.13793 · cs.LG · submitted Mar 20, 2024 · updated Apr 5, 2024
abstract · pdf · html

add comment on HN