SOTAVerified

Sound Event Localization and Detection

Given multichannel audio input, a sound event detection and localization (SELD) system outputs a temporal activation track for each of the target sound classes, along with one or more corresponding spatial trajectories when the track indicates activity. This results in a spatio-temporal characterization of the acoustic scene that can be used in a wide range of machine cognition tasks, such as inference on the type of environment, self-localization, navigation without visual input or with occluded targets, tracking of specific types of sound sources, smart-home applications, scene visualization systems, and audio surveillance, among others.

Papers

Showing 51–65 of 65 papers

TitleStatusHype
Learning Spatially-Aware Language and Audio Embeddings—0
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation—0
6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human—0
META-SELD: Meta-Learning for Fast Adaptation to the new environment in Sound Event Localization and Detection—0
Mobile Microphone Array Speech Detection and Localization in Diverse Everyday Environments—0
Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality—0
SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation—0
Sound Event Localization based on Sound Intensity Vector Refined By DNN-Based Denoising and Source Separation—0
Sound source detection, localization and classification using consecutive ensemble of CRNN models—0
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos—0
Squeeze-and-Excite ResNet-Conformers for Sound Event Localization, Detection, and Distance Estimation for DCASE 2024 Challenge—0
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling—0
SwG-former: A Sliding-Window Graph Convolutional Network for Simultaneous Spatial-Temporal Information Extraction in Sound Event Localization and Detection—0
TASK3 DCASE2021 Challenge: Sound event localization and detection using squeeze-excitation residual CNNs—0
Text-Queried Target Sound Event Localization—0
Show:102550
← PrevPage 3 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AVC-FillerNetevent-based F1 score92.8—Unverified
2VC-FillerNetevent-based F1 score71—Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (MIC)Class-dependent localization error32.2—Unverified
2Baseline (FOA)Class-dependent localization error29.3—Unverified
#ModelMetricClaimedVerifiedStatus
1DualQSELD-TCN (parallel)SELD score0.32—Unverified
#ModelMetricClaimedVerifiedStatus
1STL-SNNaccuracy98.4—Unverified
#ModelMetricClaimedVerifiedStatus
1SALSA-FOAER≤20°0.38—Unverified