about
Provably safe systems: the only path to controllable AGI (arxiv.org)
2 points by rbanffy on Sep 7, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Build powerful AI so it is proven to follow rules humans write, using AI to check the code and inspect its inner workings. This approach is argued to be the only one that guarantees safe, controlled AGI, and to become technically possible soon.

Abstract

We describe a path to humanity safely thriving with powerful Artificial General Intelligences (AGIs) by building them to provably satisfy human-specified requirements. We argue that this will soon be technically feasible using advanced AI for formal verification and mechanistic interpretability. We further argue that it is the only path which guarantees safe controlled AGI. We end with a list of challenge problems whose solution would contribute to this positive outcome and invite readers to join in this work.

Max Tegmark, Steve Omohundro
arXiv:2309.01933 · cs.CY, cs.AI, cs.LG · submitted Sep 5, 2023
abstract · pdf · html · 17 pages

add comment on HN
Also discussed: Sep 2023 (45 points, 64 comments)

" We argue that this will soon be technically feasible using advanced AI for formal verification and mechanistic interpretability."

But could AGI get in and corrupt the verification process?