In plain words: A survey sorts the recent flood of large generative AI models by what they turn into what, from text to images, video, audio, 3D shapes, code and science. It shows the field reaches far beyond chatbots like ChatGPT, touching many jobs and industries.
Abstract · ChatGPT is not all you need. A State of the Art Review of large Generative AI models
During the last two years there has been a plethora of large generative models such as ChatGPT or Stable Diffusion that have been published. Concretely, these models are able to perform tasks such as being a general question and answering system or automatically creating artistic images that are revolutionizing several sectors. Consequently, the implications that these generative models have in the industry and society are enormous, as several job positions may be transformed. For example, Generative AI is capable of transforming effectively and creatively texts to images, like the DALLE-2 model; text to 3D images, like the Dreamfusion model; images to text, like the Flamingo model; texts to video, like the Phenaki model; texts to audio, like the AudioLM model; texts to other texts, like ChatGPT; texts to code, like the Codex model; texts to scientific texts, like the Galactica model or even create algorithms like AlphaTensor. This work consists on an attempt to describe in a concise way the main models are sectors that are affected by generative AI and to provide a taxonomy of the main generative models published recently.
Roberto Gozalo-Brizuela, Eduardo C. Garrido-Merchan
arXiv:2301.04655 · cs.LG, cs.AI · submitted Jan 11, 2023
abstract · pdf · html · 22 pages
Also, what is "SOTA" for review? There isn't exactly a benchmark to compare...
In terms of comprehensiveness they don't mention PaLM + variants, which probably should be mentioned as it is currently the largest LLM with SOTA on several benchmarks (e.g. MedQA-USMLE).
In terms of correctness, I admittedly skipped to the sections I'm familiar with (LLMs) but I don't understand why they are distinguishing 'text-science' from 'text-text', they're both text-text and there is no reason why you can't, for example, adapt GPT3.5 to a scientific domain domain (some people even argue this is a better approach). A lot of powerful language models in the biomedical domain were initialized from general language models and use out-of-domain tokenizers/vocabularies (e.g. BioBERT).
The authors also make this statement regarding Galactica:
"The main advantage of [Galactica] is the ability to train on it for multiple epochs without overfitting"
This is not a unique feature of Galactica and has been done before. You're allowed to train LLMs for more than 1 epoch and in fact it can be very beneficial (see BioBERT as an example of increasing training length).
People GENERALLY don't do this because the corpus used during self-supervised training is filled with garbage/noise, so the model starts to fit to that instead of what you desire. There is nothing special about Galactica's architecture that specifically allows/encourages longer training cycles but rather they curated the dataset to minimize garbage. As another example, my research involves radiology NLP and when doing domain adaptive pretraining on a highly curated dataset we have been going up to 8 epochs without overfitting.