SOTAVerified

Audio captioning

Audio Captioning is the task of describing audio using text. The general approach is to use an audio encoder to encode the audio (example: PANN, CAV-MAE), and to use a decoder (example: transformer) to generate the text. To judge the quality of audio captions, though machine translation metrics (BLEU, METEOR, ROUGE) and image captioning metrics (SPICE, CIDER) are used, they are not very well-suited. Attempts have been made to use pretrained language model based metrics such as Sentence-BERT.

Papers

Showing 1–10 of 119 papers

TitleStatusHype
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue AbilitiesCode5
Improving Text-To-Audio Models with Synthetic CaptionsCode5
LLMs can see and hear without any trainingCode3
SALMONN: Towards Generic Hearing Abilities for Large Language ModelsCode3
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language ModelsCode3
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPTCode2
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual FusionCode2
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio CaptioningCode2
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language ModelsCode2
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning PerformanceCode2
Show:102550
← PrevPage 1 of 12Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1VASTCIDEr0.78—Unverified
2VALORCIDEr0.74—Unverified
3MQ-CapSPIDEr0.52—Unverified
4SLAM-AACSPIDEr0.52—Unverified
5LAVCapSPIDEr0.52—Unverified
6EnCLAP++-largeSPIDEr0.51—Unverified
7AutoCapSPIDEr0.51—Unverified
8LOAESPIDEr0.51—Unverified
9EnCLAP++-baseSPIDEr0.5—Unverified
10EnCLAP-largeSPIDEr0.5—Unverified
#ModelMetricClaimedVerifiedStatus
1VASTCIDEr0.52—Unverified
2VALORCIDEr0.42—Unverified
3SLAM-AACSPIDEr0.33—Unverified
4LOAESPIDEr0.33—Unverified
5MQ-CapSPIDEr0.32—Unverified
6EnsembleSPIDEr0.32—Unverified
7Audio Flamingo (Pengi trainset)SPIDEr0.31—Unverified
8Ensemble-RLSPIDEr0.3—Unverified
9Qwen-AudioSPIDEr0.29—Unverified
10EnsembleSPIDEr0.21—Unverified