about
AudioPaLM: A large language model that can speak and listen (arxiv.org)
69 points by sheepscreek on Jun 26, 2023 | hide | past | pdf | 11 comments on HN

In plain words: One system merges a text language model with a speech model, so it handles both text and speech while keeping the speaker's tone and the text model's knowledge. It beat the best speech translation systems, even on language pairs it never saw in training.

Abstract · AudioPaLM: A Large Language Model That Can Speak and Listen

We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a unified multimodal architecture that can process and generate text and speech with applications including speech recognition and speech-to-speech translation. AudioPaLM inherits the capability to preserve paralinguistic information such as speaker identity and intonation from AudioLM and the linguistic knowledge present only in text large language models such as PaLM-2. We demonstrate that initializing AudioPaLM with the weights of a text-only large language model improves speech processing, successfully leveraging the larger quantity of text training data used in pretraining to assist with the speech tasks. The resulting model significantly outperforms existing systems for speech translation tasks and has the ability to perform zero-shot speech-to-text translation for many languages for which input/target language combinations were not seen in training. AudioPaLM also demonstrates features of audio language models, such as transferring a voice across languages based on a short spoken prompt. We release examples of our method at https://google-research.github.io/seanet/audiopalm/examples

Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, et al.
arXiv:2306.12925 · cs.CL, cs.AI, cs.SD, eess.AS, stat.ML · submitted Jun 22, 2023
abstract · pdf · html · Technical report

add comment on HN

Preserving the speakers voice after translation is super cool. It’s one of the things you don’t really think about, but voice inflection and identity is missing with our current translation tools.

I like to watch political speeches from non english speaking politicians, and the speakers tone can easily be lost in translation. Emphasis is hard to discern when you don’t know which spoken word maps to which word in the subtitles. Dubbed speeches are even worse in that respect.

Uh oh, here come the multimedia multimodels.

I see a lot of talk about transformers LLMs being close to "topping out," which I am skeptical of for many reasons, but not the least of which is prompting/outputs other than pure text.

Is there really a difference between a multimodal vs a text LLM + stable diffusion?
Multimodal can refer to a lot of different types of models, but feeding LLM text into stable diffusion definitely doesn’t count.

LLaVA is the first one that comes to my mind, it takes images and text as input and outputs text.

There’s an unreleased version of GPT4 that can do that same thing.

Sure technically not the same, but won't there be the same affect?

How do our brains work? Isn't there a separation between image and text processing?

Surely there needs to be some amount of training with both models in the loop before it can be considered a multimodal system.
What do you mean?
That if you have 3 separate models, text -> text, image -> text and text -> image you can just glue them together to make it behave like a multimodal.

(Just like gpt4 is rumored to be a few different sub models and not just one giant model)

Additional discussion from post two days ago: https://news.ycombinator.com/item?id=36443676
Demo website with speech-speech translation examples https://google-research.github.io/seanet/audiopalm/examples/
Absolutely amazing. Anything like this available to play with?