SOTAVerified

Keyword Spotting

In speech processing, keyword spotting deals with the identification of keywords in utterances.

( Image credit: Simon Grest )

Papers

Showing 251–300 of 407 papers

TitleStatusHype
Weight-importance sparse training in keyword spotting—0
Word Searching in Scene Image and Video Frame in Multi-Script Scenario using Dynamic Shape Coding—0
Work in Progress: Linear Transformers for TinyML—0
WSRNet: Joint Spotting and Recognition of Handwritten Words—0
Harnessing the Power of Explanations for Incremental Training: A LIME-Based Approach—0
Zero-Shot Federated Learning with New Classes for Audio Classification—0
Zero-Shot Temporal Resolution Domain Adaptation for Spiking Neural Networks—0
0/1 Deep Neural Networks via Block Coordinate Descent—0
A Lightweight dynamic filter for keyword spotting—0
LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting—0
Locale Encoding For Scalable Multilingual Keyword Spotting Models—0
Low-bit quantization and quantization-aware training for small-footprint keyword spotting—0
Low-Power Low-Latency Keyword Spotting and Adaptive Control with a SpiNNaker 2 Prototype and Comparison with Loihi—0
Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings—0
Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input—0
Matching Latent Encoding for Audio-Text based Keyword Spotting—0
Maximum-Entropy Adversarial Audio Augmentation for Keyword Spotting—0
Max-Pooling Loss Training of Long Short-Term Memory Networks for Small-Footprint Keyword Spotting—0
Meta-Learning for improving rare word recognition in end-to-end ASR—0
Metric Learning for Keyword Spotting—0
Metric Learning for User-defined Keyword Spotting—0
Micro-power spoken keyword spotting on Xylo Audio 2—0
Modular approach to data preprocessing in ALOHA and application to a smart industry use case—0
More than words: Advancements and challenges in speech recognition for singing—0
Morphological Segmentation for Keyword Spotting—0
Multi-layer Attention Mechanism for Speech Keyword Recognition—0
Multilingual acoustic word embeddings for zero-resource languages—0
Multilingual Query-by-Example Keyword Spotting with Metric Learning and Phoneme-to-Embedding Mapping—0
Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis—0
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio—0
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting—0
Multitaper mel-spectrograms for keyword spotting—0
Multi-task Learning with Cross Attention for Keyword Spotting—0
Multi-Task Network for Noise-Robust Keyword Spotting and Speaker Verification using CTC-based Soft VAD and Global Query Attention—0
Multi-task Voice Activated Framework using Self-supervised Learning—0
Domain Aware Training for Far-field Small-footprint Keyword Spotting—0
Neural Architecture Search For Keyword Spotting—0
Neural Morphological Analysis: Encoding-Decoding Canonical Segments—0
Neural Networks for Keyword Spotting on IoT Devices—0
Noise-Agnostic Multitask Whisper Training for Reducing False Alarm Errors in Call-for-Help Detection—0
Noise-Robust Hearing Aid Voice Control—0
Noisy student-teacher training for robust keyword spotting—0
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting—0
NTU System at MediaEval 2015: Zero Resource Query by Example Spoken Term Detection Using Deep and Recurrent Neural Networks—0
On-Device Constrained Self-Supervised Speech Representation Learning for Keyword Spotting via Knowledge Distillation—0
On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems—0
On evaluating CNN representations for low resource medical image classification—0
Online Keyword Spotting with a Character-Level Recurrent Neural Network—0
On the Efficiency of Integrating Self-supervised Learning and Meta-learning for User-defined Few-shot Keyword Spotting—0
On the Non-Associativity of Analog Computations—0
Show:102550
← PrevPage 6 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NNI non-filtered(for the development set)Cnxe6.09—Unverified
2NNI Choi(for the development set)Cnxe5.89—Unverified
3NTU rnn (eval)Cnxe2.01—Unverified
4NTU dtw (eval)Cnxe2.01—Unverified
5NTU dtw (dev)Cnxe2.01—Unverified
6NTU rnn (dev)Cnxe2.01—Unverified
7ELiRF SDTW (eval)Cnxe1.19—Unverified
8ELiRF SDTW-avg (eval)Cnxe1.07—Unverified
9ELiRF SDTW (dev)Cnxe1.07—Unverified
10CUNY [Subseq+MFCC] (eval)Cnxe1.07—Unverified
#ModelMetricClaimedVerifiedStatus
1WaveFormerGoogle Speech Commands V2 1298.8—Unverified
2QNNGoogle Speech Commands V2 3598.6—Unverified
3TripletLoss-res15Google Speech Commands V1 1298.56—Unverified
4M2DGoogle Speech Commands V2 3598.5—Unverified
5EAT-SGoogle Speech Commands V2 3598.15—Unverified
6Audio Spectrogram TransformerGoogle Speech Commands V2 3598.11—Unverified
7EdgeCRNN 2.0×Google Speech Commands V2 1298.05—Unverified
8BC-ResNet-8Google Speech Commands V1 1298—Unverified
9HTS-ATGoogle Speech Commands V2 3598—Unverified
10Wav2KWSGoogle Speech Commands V1 1297.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Stacked 1D CNNError Rate1.99—Unverified
2End-to-end DNN-HMMError Rate1.7—Unverified
3HEiMDaLError Rate0.45—Unverified
#ModelMetricClaimedVerifiedStatus
1Res26Accuracy95.88—Unverified
2EfficientNet-A0 + SA + TLAccuracy95.83—Unverified
#ModelMetricClaimedVerifiedStatus
1QuaternionNeuralNetworkAccuracy (10-fold)98.53—Unverified
2SSAMBAAccuracy (10-fold)97.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TensorFlow's model version 2TFMA89.7—Unverified
2TensorFlow's model version 1TFMA85.4—Unverified
#ModelMetricClaimedVerifiedStatus
12D-ConvNetAccuracy (%)95.4—Unverified
21D-ConvNetAccuracy (%)93.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Quaternion Neural NetworksAccuracy(10-fold)98.53—Unverified
#ModelMetricClaimedVerifiedStatus
1MicroNet-KWS-LAccuracy95.3—Unverified