about
Nvidia Audio Flamingo, Audio LM with Few-Shot Learning and Dialogue Abilities (arxiv.org)
1 point by alok-g on Feb 14, 2024 | hide | past | pdf | discuss on HN

In plain words: A language model that listens to audio, including everyday sounds and tone of voice, and can pick up new tasks from just a few examples while chatting over several turns. It beat the best previous audio-understanding systems across a wide range of listening tests.

Abstract · Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs. In this paper, we propose Audio Flamingo, a novel audio language model with 1) strong audio understanding abilities, 2) the ability to quickly adapt to unseen tasks via in-context learning and retrieval, and 3) strong multi-turn dialogue abilities. We introduce a series of training techniques, architecture design, and data strategies to enhance our model with these abilities. Extensive evaluations across various audio understanding tasks confirm the efficacy of our method, setting new state-of-the-art benchmarks. Our demo website is https://audioflamingo.github.io/ and the code is open-sourced at https://github.com/NVIDIA/audio-flamingo.

Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, Bryan Catanzaro
arXiv:2402.01831 · cs.SD, cs.LG, eess.AS · submitted Feb 2, 2024 · updated May 28, 2024
abstract · pdf · html · ICML 2024

add comment on HN