SOTAVerified

Keyword Spotting

In speech processing, keyword spotting deals with the identification of keywords in utterances.

( Image credit: Simon Grest )

Papers

Showing 351–400 of 407 papers

TitleStatusHype
Spoken Language Identification using ConvNets—0
Spot keywords from very noisy and mixed speech—0
ST-KeyS: Self-Supervised Transformer for Keyword Spotting in Historical Handwritten Documents—0
Streaming Small-Footprint Keyword Spotting using Sequence-to-Sequence Models—0
Let SSMs be ConvNets: State-space Modeling with Optimal Tensor ContractionsCode0
Depth Pruning with Auxiliary Networks for TinyMLCode0
Keyword Spotting Simplified: A Segmentation-Free Approach using Character Counting and CTC re-scoringCode0
Open-source FPGA-ML codesign for the MLPerf Tiny BenchmarkCode0
Honkling: In-Browser Personalization for Ubiquitous Keyword SpottingCode0
Cluster-based pruning techniques for audio dataCode0
Avoid Overfitting User Specific Information in Federated Keyword SpottingCode0
End-to-end Keyword Spotting using Xception-1dCode0
Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword SpottingCode0
Keyword localisation in untranscribed speech using visually grounded speech modelsCode0
Efficient keyword spotting using dilated convolutions and gatingCode0
Stochastic Adaptive Neural Architecture Search for Keyword SpottingCode0
Semi-Supervised Federated Learning for Keyword SpottingCode0
An Investigation of Few-Shot Learning in Spoken Term ClassificationCode0
Building and benchmarking an Arabic Speech Commands dataset for small-footprint keyword spottingCode0
Efficient Keyword Spotting by capturing long-range interactions with Temporal Lambda NetworksCode0
What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy ConditionsCode0
JavaScript Convolutional Neural Networks for Keyword Spotting in the Browser: An Experimental AnalysisCode0
Integrated Parameter-Efficient Tuning for General-Purpose Audio ModelsCode0
Hello Edge: Keyword Spotting on MicrocontrollersCode0
Indian EmoSpeech Command Dataset: A dataset for emotion based speech recognition in the wildCode0
ImportantAug: a data augmentation agent for speechCode0
Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-ConvolutionsCode0
GTM-UVigo Systems for the Query-by-Example Search on Speech Task at MediaEval 2015Code0
Unsupervised Speech Representation Pooling Using Vector QuantizationCode0
AraSpot: Arabic Spoken Command SpottingCode0
ed-cec: improving rare word recognition using asr postprocessing based on error detection and context-aware error correctionCode0
What’s Cookin’? Interpreting Cooking Videos using Text, Speech and VisionCode0
What's Cookin'? Interpreting Cooking Videos using Text, Speech and VisionCode0
DONUT: CTC-based Query-by-Example Keyword SpottingCode0
Filler Word Detection and Classification: A Dataset and BenchmarkCode0
Learning Delays Through Gradients and Structure: Emergence of Spatiotemporal Patterns in Spiking Neural NetworksCode0
Audiomer: A Convolutional Transformer For Keyword SpottingCode0
Audio Explanation Synthesis with Generative Foundation ModelsCode0
Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource DevicesCode0
Boosting keyword spotting through on-device learnable user speech characteristicsCode0
Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and AlgorithmsCode0
TACos: Learning Temporally Structured Embeddings for Few-Shot Keyword Spotting with Dynamic Time WarpingCode0
Temporal Convolution for Real-time Keyword Spotting on Mobile DevicesCode0
Neural ODE with Temporal Convolution and Time Delay Neural Networks for Small-Footprint Keyword SpottingCode0
Neuromorphic Keyword Spotting with Pulse Density Modulation MEMS MicrophonesCode0
Federated Learning for Keyword SpottingCode0
READ-BAD: A New Dataset and Evaluation Scheme for Baseline Detection in Archival DocumentsCode0
Noise-Robust Keyword Spotting through Self-supervised PretrainingCode0
Temporal Feedback Convolutional Recurrent Neural Networks for Speech Command RecognitionCode0
Trainable Frontend For Robust and Far-Field Keyword SpottingCode0
Show:102550
← PrevPage 8 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NNI non-filtered(for the development set)Cnxe6.09—Unverified
2NNI Choi(for the development set)Cnxe5.89—Unverified
3NTU rnn (eval)Cnxe2.01—Unverified
4NTU dtw (eval)Cnxe2.01—Unverified
5NTU dtw (dev)Cnxe2.01—Unverified
6NTU rnn (dev)Cnxe2.01—Unverified
7ELiRF SDTW (eval)Cnxe1.19—Unverified
8ELiRF SDTW-avg (eval)Cnxe1.07—Unverified
9ELiRF SDTW (dev)Cnxe1.07—Unverified
10CUNY [Subseq+MFCC] (eval)Cnxe1.07—Unverified
#ModelMetricClaimedVerifiedStatus
1WaveFormerGoogle Speech Commands V2 1298.8—Unverified
2QNNGoogle Speech Commands V2 3598.6—Unverified
3TripletLoss-res15Google Speech Commands V1 1298.56—Unverified
4M2DGoogle Speech Commands V2 3598.5—Unverified
5EAT-SGoogle Speech Commands V2 3598.15—Unverified
6Audio Spectrogram TransformerGoogle Speech Commands V2 3598.11—Unverified
7EdgeCRNN 2.0×Google Speech Commands V2 1298.05—Unverified
8BC-ResNet-8Google Speech Commands V1 1298—Unverified
9HTS-ATGoogle Speech Commands V2 3598—Unverified
10Wav2KWSGoogle Speech Commands V1 1297.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Stacked 1D CNNError Rate1.99—Unverified
2End-to-end DNN-HMMError Rate1.7—Unverified
3HEiMDaLError Rate0.45—Unverified
#ModelMetricClaimedVerifiedStatus
1Res26Accuracy95.88—Unverified
2EfficientNet-A0 + SA + TLAccuracy95.83—Unverified
#ModelMetricClaimedVerifiedStatus
1QuaternionNeuralNetworkAccuracy (10-fold)98.53—Unverified
2SSAMBAAccuracy (10-fold)97.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TensorFlow's model version 2TFMA89.7—Unverified
2TensorFlow's model version 1TFMA85.4—Unverified
#ModelMetricClaimedVerifiedStatus
12D-ConvNetAccuracy (%)95.4—Unverified
21D-ConvNetAccuracy (%)93.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Quaternion Neural NetworksAccuracy(10-fold)98.53—Unverified
#ModelMetricClaimedVerifiedStatus
1MicroNet-KWS-LAccuracy95.3—Unverified