Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

2024-10-14Unverified0· sign in to hype

Choi Changin, Lim Sungjun, Rhee Wonjong

Unverified — Be the first to reproduce this paper.

Abstract

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose Generation-Assisted Multimodal Querying, which generates a text description of the input audio to enable multimodal querying. This approach aligns the query modality with the audio-text structure of the knowledge base, leading to more effective retrieval. Furthermore, we introduce a novel progressive learning strategy that gradually increases the number of interleaved audio-text pairs to enhance the training process. Our experiments on AudioCaps, Clotho, and Auto-ACD demonstrate that our approach achieves state-of-the-art results across these benchmarks.

Tasks

AudioCaps Audio captioning RAG Retrieval Retrieval-augmented Generation

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
AudioCaps	MQ-Cap	SPIDEr	0.52	—	Unverified
Clotho	MQ-Cap	SPIDEr	0.32	—	Unverified

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

Abstract

Tasks

Benchmark Results

Reproductions