Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

2024-02-02Code Available5· sign in to hype

Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, Bryan Catanzaro

Code Available — Be the first to reproduce this paper.

Code

github.com/NVIDIA/audio-flamingo
OfficialIn paperpytorch★ 1,020

Abstract

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs. In this paper, we propose Audio Flamingo, a novel audio language model with 1) strong audio understanding abilities, 2) the ability to quickly adapt to unseen tasks via in-context learning and retrieval, and 3) strong multi-turn dialogue abilities. We introduce a series of training techniques, architecture design, and data strategies to enhance our model with these abilities. Extensive evaluations across various audio understanding tasks confirm the efficacy of our method, setting new state-of-the-art benchmarks. Our demo website is https://audioflamingo.github.io/ and the code is open-sourced at https://github.com/NVIDIA/audio-flamingo.

Tasks

Acoustic Scene Classification Audio captioning Few-Shot Learning In-Context Learning Language Modeling Language Modelling Retrieval Retrieval-augmented Few-shot In-context Audio Captioning Zero-shot Audio Captioning

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
CochlScene	Audio Flamingo	1:1 Accuracy	0.83	—	Unverified

Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Code

Abstract

Tasks

Benchmark Results

Reproductions