about
On the Opportunities and Risks of Foundation Models (arxiv.org)
20 points by satorii on Aug 23, 2021 | hide | past | pdf | 6 comments on HN

In plain words: Huge AI systems trained on broad data can be adapted to many tasks, from language to healthcare. Their scale brings surprising new abilities but also spreads one model's flaws everywhere, and we still don't understand how they work or fail.

Abstract

AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. We call these models foundation models to underscore their critically central yet incomplete character. This report provides a thorough account of the opportunities and risks of foundation models, ranging from their capabilities (e.g., language, vision, robotics, reasoning, human interaction) and technical principles(e.g., model architectures, training procedures, data, systems, security, evaluation, theory) to their applications (e.g., law, healthcare, education) and societal impact (e.g., inequity, misuse, economic and environmental impact, legal and ethical considerations). Though foundation models are based on standard deep learning and transfer learning, their scale results in new emergent capabilities,and their effectiveness across so many tasks incentivizes homogenization. Homogenization provides powerful leverage but demands caution, as the defects of the foundation model are inherited by all the adapted models downstream. Despite the impending widespread deployment of foundation models, we currently lack a clear understanding of how they work, when they fail, and what they are even capable of due to their emergent properties. To tackle these questions, we believe much of the critical research on foundation models will require deep interdisciplinary collaboration commensurate with their fundamentally sociotechnical nature.

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, et al.
arXiv:2108.07258 · cs.LG, cs.AI, cs.CY · submitted Aug 16, 2021 · updated Jul 12, 2022
abstract · pdf · html · Authored by the Center for Research on Foundation Models (CRFM) at the Stanford Institute for Human-Centered Artificial Intelligence (HAI). Report page with citation guidelines: https://crfm.stanford.edu/report.html

add comment on HN
Also discussed: Jul 2024 (1 point, 0 comments) · Mar 2023 (1 point, 0 comments) · Aug 2021 (2 points, 1 comment) · Aug 2021 (2 points, 0 comments) · Aug 2021 (3 points, 1 comment) · Aug 2021 (4 points, 0 comments)

I'd suggest using the original article title. [Edit, it's been updated]

Still have to read the article. It's great to see people exploring this. From the first "language models are unsupervised multitask learners" type papers, i wish there had been more emphasis that the various behaviors these models have are essentially a side effect of learning some kind of self supervision task. A model has been trained to e.g. predict the next word given previous words, and we're happy to discover that it can be repurposed as a chatbot. And then people find the chatbot has some undesirable behaviors, and talk about fairness and governance and all that. When the basic point is the model was never really trained to do any of that, its just a word predictor. Why did you ever think it would be OK to just let it run wild on some other task?

All that to say, a big problem in AI/ML is models getting used for things they have no business being used for, and them people being at best underwhelmed, or harmed or offended by the results. The first step should be asking why is this model suitable for making the prediction I'm asking it to, and I think closer scrutiny on what these "foundation models" actually do is a good direction.

(Title changed now. Submitted title was "What is this new AI term, foundation models".)
To answer the question the original poster apparently had, here are the first two sentences of the abstract:

> AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. We call these models foundation models to underscore their critically central yet incomplete character.

I'm curious about the format/formatting of this paper. There are a few visual roadmaps to the various sections and subsections throughout the paper, complete with drawings/iconography (clip art?). I haven't seen anything like this before in an academic paper. Is it something that's becoming popular in certain research communities?
Interesting findings! Not sure about whether it is within certain research communities or a broader trend.

But Clip cloud be a good plug-in for nowadays writings/design then, something like Clip empowered unsplash.

It's a 212 page scientific report, not a traditional article.

Could the vision system not be the AI foundational model? Car vision, specifically Tesla state-of-the-art is not mentioned. I see bit of NLP bias.