Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

2023-10-18Unverified0· sign in to hype

Yiyang Su, Ali Vosoughi, Shijian Deng, Yapeng Tian, Chenliang Xu

Unverified — Be the first to reproduce this paper.

Abstract

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduces a novel "Audio-Visual Scene-Aware Separation" (AVSA-Sep) framework. It includes a semantic parser for visible and invisible sounds and a separator for scene-informed separation. AVSA-Sep successfully separates both sound types, with joint training and cross-modal alignment enhancing effectiveness.

Tasks

cross-modal alignment

Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

Abstract

Tasks

Reproductions