SOTAVerified

Sound Event Detection

Sound Event Detection (SED) is the task of recognizing the sound events and their respective temporal start and end time in a recording. Sound events in real life do not always occur in isolation, but tend to considerably overlap with each other. Recognizing such overlapping sound events is referred as polyphonic SED.

Source: A report on sound event detection with different binaural features

Papers

Showing 151–194 of 194 papers

TitleStatusHype
Improving Weakly Supervised Sound Event Detection with Causal Intervention—0
Incremental Learning Algorithm for Sound Event Detection—0
Interactive Dual-Conformer with Scene-Inspired Mask for Soft Sound Event Detection—0
Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations—0
Leveraging Audio-Tagging Assisted Sound Event Detection using Weakified Strong Labels and Frequency Dynamic Convolutions—0
Lightweight Sound Event Detection Model with RepVGG Architecture—0
Mixstyle based Domain Generalization for Sound Event Detection with Heterogeneous Training Data—0
Multi-Branch Learning for Weakly-Labeled Sound Event Detection—0
Multichannel Sound Event Detection Using 3D Convolutional Neural Networks for Learning Inter-channel Features—0
Multi-dimensional frequency dynamic convolution with confident mean teacher for sound event detection—0
Multi-encoder attention-based architectures for sound recognition with partial visual assistance—0
Multitask frame-level learning for few-shot sound event detection—0
Multitask vocal burst modeling with ResNets and pre-trained paralinguistic Conformers—0
Nonverbal Sound Detection for Disordered Speech—0
Online Active Learning For Sound Event Detection—0
Optimizing Temporal Resolution Of Convolutional Recurrent Neural Networks For Sound Event Detection—0
Peer Collaborative Learning for Polyphonic Sound Event Detection—0
Power pooling: An adaptive pooling function for weakly labelled sound event detection—0
Proposal-based Few-shot Sound Event Detection for Speech and Environmental Sounds with Perceivers—0
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection—0
Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events—0
RCRNN-based Sound Event Detection System with Specific Speech Resolution—0
LOCUS: LOcalization with Channel Uncertainty and Sporadic Energy—0
Robust and Interpretable Temporal Convolution Network for Event Detection in Lung Sound Recordings—0
Robust detection of overlapping bioacoustic sound events—0
SEED: Sound Event Early Detection via Evidential Uncertainty—0
SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation—0
Selective Pseudo-labeling and Class-wise Discriminative Fusion for Sound Event Detection—0
Self-training with noisy student model and semi-supervised loss function for dcase 2021 challenge task 4—0
Semi-supervised Sound Event Detection with Local and Global Consistency Regularization—0
Semi-supervsied Learning-based Sound Event Detection using Freuqency Dynamic Convolution with Large Kernel Attention for DCASE Challenge 2023 Task 4—0
Soft-Median Choice: An Automatic Feature Smoothing Method for Sound Event Detection—0
SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform—0
Sound Event Detection and Localization with Distance Estimation—0
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4—0
Sound Event Detection in Domestic Environments using Dense Recurrent Neural Network—0
Sound Event Detection in Multichannel Audio Using Spatial and Harmonic Features—0
Sound Event Detection in Multichannel Audio using Convolutional Time-Frequency-Channel Squeeze and Excitation—0
Sound Event Detection in Synthetic Audio: Analysis of the DCASE 2016 Task Results—0
Sound Event Detection in Urban Audio With Single and Multi-Rate PCEN—0
Sound event detection via dilated convolutional recurrent neural networks—0
Sound Event Detection with Binary Neural Networks on Tightly Power-Constrained IoT Devices—0
SwG-former: A Sliding-Window Graph Convolutional Network for Simultaneous Spatial-Temporal Information Extraction in Sound Event Localization and Detection—0
Synthetic data enables context-aware bioacoustic sound event detection—0
Show:102550
← PrevPage 4 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ATST-SEDevent-based F1 score63.4—Unverified
2SE-CRNN-16 with DualKDevent-based F1 score55.6—Unverified
3FDY-CRNNevent-based F1 score54—Unverified
4HTS-ATevent-based F1 score50.7—Unverified
5RCTevent-based F1 score49.62—Unverified
6FiltAug SEDevent-based F1 score49.6—Unverified
7SED-SSep baseline dcase task 4 2020 v2event-based F1 score40.7—Unverified
8Baseline dcase task 4 2020 v2event-based F1 score39—Unverified
9Baselineevent-based F1 score25.8—Unverified
10MAT-SEDPSDS10.59—Unverified
#ModelMetricClaimedVerifiedStatus
1PHC SEDnet n=8Error Rate0.56—Unverified
2Quaternion SEDnetError Rate0.52—Unverified
3PHC SEDnet n=16Error Rate0.51—Unverified
4PHC SEDnet n=4Error Rate0.45—Unverified
5PHC SEDnet n=2Error Rate0.39—Unverified
#ModelMetricClaimedVerifiedStatus
1CRNN (with BEATs + Separation)PSDS1 (-5dB)0.13—Unverified
2CRNN (with BEATs)PSDS1 (-5dB)0.07—Unverified
3CRNN (WildDESED + Curriculrm learning)PSDS1 (-5dB)0.05—Unverified
4CRNN (WildDESED)PSDS1 (-5dB)0.05—Unverified
5CRNNPSDS1 (-5dB)0.02—Unverified
#ModelMetricClaimedVerifiedStatus
1DENetRank-1 Recognition Rate0.98—Unverified
#ModelMetricClaimedVerifiedStatus
1DENetRank-1 Recognition Rate1—Unverified