about
Unsolved Problems in ML Safety (arxiv.org)
61 points by pramodbiligiri on Oct 9, 2021 | hide | past | pdf | 3 comments on HN

In plain words: A roadmap sorts machine learning safety research into four tasks: resisting hazards, spotting them, making models less harmful by design, and fixing risks built into the wider system. It sharpens the field's vague worries into concrete, ready-to-study problems with suggested research directions.

Abstract

Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the technical problems that the field needs to address. We present four problems ready for research, namely withstanding hazards ("Robustness"), identifying hazards ("Monitoring"), reducing inherent model hazards ("Alignment"), and reducing systemic hazards ("Systemic Safety"). Throughout, we clarify each problem's motivation and provide concrete research directions.

Dan Hendrycks, Nicholas Carlini, John Schulman, Jacob Steinhardt
arXiv:2109.13916 · cs.LG, cs.AI, cs.CL, cs.CV · submitted Sep 28, 2021 · updated Jun 16, 2022
abstract · pdf · html · Position Paper

add comment on HN
Also discussed: Oct 2021 (1 point, 0 comments) · Sep 2021 (2 points, 0 comments)

Some overlap here with the 2017 paper from Google Brain et al: Concrete Problems in AI Safety: https://arxiv.org/pdf/1606.06565.pdf

A nice video from the wonderful AI safety communicator Robert Miles outlines those problems: https://www.youtube.com/watch?v=AjyM-f8rDpg

ML=machine learning; not meta language.
These are unsolved in the sense that there is no formal or comprehensive solution to them, and in the limit that means that ML cannot be used for some applications.

On the other hand there are mitigations that can be implemented to increase the confidence that we have in ML centric solutions - and that can bring some applications back into scope.

Figure 5 shows this using the metaphor of layers of swiss cheese that other folk have used with relation to stopping Covid. Of course we can't stop Covid completely, but by being careful and getting a vaccine, ventilation and masks you can improve your chances.