| Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention | Oct 28, 2022 | AudioCapsAudio captioning | CodeCode Available | 1 | 5 |
| Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval | Aug 21, 2024 | AudioCapsContrastive Learning | CodeCode Available | 0 | 5 |
| Weakly-supervised Automated Audio Captioning via text only training | Sep 21, 2023 | AudioCapsAudio captioning | CodeCode Available | 0 | 5 |
| SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs | Oct 12, 2024 | AudioCapsAudio captioning | CodeCode Available | 0 | 5 |
| ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors | Feb 20, 2025 | AudioCapsContrastive Learning | CodeCode Available | 0 | 5 |
| MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation | Jun 15, 2024 | AudioCapsImage Generation | CodeCode Available | 0 | 5 |
| Accommodating Audio Modality in CLIP for Multimodal Processing | Mar 12, 2023 | AudioCapsContrastive Learning | CodeCode Available | 0 | 5 |
| AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS | Nov 15, 2021 | AudioCapsAudio captioning | CodeCode Available | 0 | 5 |
| Audiobox: Unified Audio Generation with Natural Language Prompts | Dec 25, 2023 | AudioCapsAudio Generation | —Unverified | 0 | 0 |
| IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling | May 31, 2025 | AudioCapsAudio Generation | —Unverified | 0 | 0 |
| Joint Speech Recognition and Audio Captioning | Feb 3, 2022 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Killing two birds with one stone: Can an audio captioning system also be used for audio-text retrieval? | Aug 29, 2023 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Language-based Audio Retrieval with Co-Attention Networks | Dec 30, 2024 | AudioCapsLearning Semantic Representations | —Unverified | 0 | 0 |
| TAIL: Text-Audio Incremental Learning | Mar 6, 2025 | AudioCapsIncremental Learning | —Unverified | 0 | 0 |
| Leveraging Pre-trained BERT for Audio Captioning | Mar 6, 2022 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning | May 28, 2025 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval | Mar 15, 2024 | AudioCapsContrastive Learning | —Unverified | 0 | 0 |
| VoiceLDM: Text-to-Speech with Environmental Context | Sep 24, 2023 | AudioCapstext-to-speech | —Unverified | 0 | 0 |
| Text-to-Audio Generation Synchronized with Videos | Mar 8, 2024 | AudioCapsAudio Generation | —Unverified | 0 | 0 |
| Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model | Mar 12, 2025 | AudioCapsContrastive Learning | —Unverified | 0 | 0 |
| AC/DC: LLM-based Audio Comprehension via Dialogue Continuation | Jun 12, 2025 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Rethinking Transfer and Auxiliary Learning for Improving Audio Captioning Transformer | Aug 20, 2023 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Retrieval-Augmented Text-to-Audio Generation | Sep 14, 2023 | AudioCapsAudio Generation | —Unverified | 0 | 0 |
| Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning | Feb 8, 2025 | AudioCapsAudio captioning | —Unverified | 0 | 0 |
| Audio-text Retrieval in Context | Mar 25, 2022 | AudioCapsRetrieval | —Unverified | 0 | 0 |