about
On the Opportunities and Risks of Foundation Models (arxiv.org)
3 points by Anon84 on Aug 18, 2021 | hide | past | pdf | 1 comment on HN

In plain words: Huge AI systems trained on broad data can be adapted to many tasks, from language to healthcare. Their scale brings surprising new abilities but also spreads one model's flaws everywhere, and we still don't understand how they work or fail.

Abstract

AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. We call these models foundation models to underscore their critically central yet incomplete character. This report provides a thorough account of the opportunities and risks of foundation models, ranging from their capabilities (e.g., language, vision, robotics, reasoning, human interaction) and technical principles(e.g., model architectures, training procedures, data, systems, security, evaluation, theory) to their applications (e.g., law, healthcare, education) and societal impact (e.g., inequity, misuse, economic and environmental impact, legal and ethical considerations). Though foundation models are based on standard deep learning and transfer learning, their scale results in new emergent capabilities,and their effectiveness across so many tasks incentivizes homogenization. Homogenization provides powerful leverage but demands caution, as the defects of the foundation model are inherited by all the adapted models downstream. Despite the impending widespread deployment of foundation models, we currently lack a clear understanding of how they work, when they fail, and what they are even capable of due to their emergent properties. To tackle these questions, we believe much of the critical research on foundation models will require deep interdisciplinary collaboration commensurate with their fundamentally sociotechnical nature.

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, et al.
arXiv:2108.07258 · cs.LG, cs.AI, cs.CY · submitted Aug 16, 2021 · updated Jul 12, 2022
abstract · pdf · html · Authored by the Center for Research on Foundation Models (CRFM) at the Stanford Institute for Human-Centered Artificial Intelligence (HAI). Report page with citation guidelines: https://crfm.stanford.edu/report.html

add comment on HN
Also discussed: Jul 2024 (1 point, 0 comments) · Mar 2023 (1 point, 0 comments) · Aug 2021 (20 points, 6 comments) · Aug 2021 (2 points, 1 comment) · Aug 2021 (2 points, 0 comments) · Aug 2021 (4 points, 0 comments)

I expected a more in-depth technical paper with a clear project vision, but what I saw was basically Reframing Superintelligence Comprehensive AI Services as General Intelligence (2019) by K. Eric Drexler: https://www.fhi.ox.ac.uk/reframing/ ... with no citation or otherwise credit given to this prior work. In the post-GPT3 era, discussing this path of technological development is as straightforward as trivial extrapolation. Expecting things to develop this way in the beginning of 2019, before GPT2, would require prescient analysis.