SOTAVerified

Keyword Spotting

In speech processing, keyword spotting deals with the identification of keywords in utterances.

( Image credit: Simon Grest )

Papers

Showing 301–350 of 407 papers

TitleStatusHype
Open-vocabulary Keyword-spotting with Adaptive Instance Normalization—0
Optimize what matters: Training DNN-HMM Keyword Spotting Model Using End Metric—0
Orthogonality Constrained Multi-Head Attention For Keyword Spotting—0
PATE-AAE: Incorporating Adversarial Autoencoder into Private Aggregation of Teacher Ensembles for Spoken Command Classification—0
PBSM: Backdoor attack against Keyword spotting based on pitch boosting and sound masking—0
Performance-Oriented Neural Architecture Search—0
Personalized Keyword Spotting through Multi-task Learning—0
Personalizing Keyword Spotting with Speaker Information—0
Phone Based Keyword Spotting for Transcribing Very Low Resource Languages—0
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment—0
Plug-and-Play Multilingual Few-shot Spoken Words Recognition—0
Polish Read Speech Corpus for Speech Tools and Services—0
PEAF: Learnable Power Efficient Analog Acoustic Features for Audio Recognition—0
PQK: Model Compression via Pruning, Quantization, and Knowledge Distillation—0
Predicting detection filters for small footprint open-vocabulary keyword spotting—0
Production federated keyword spotting via distillation, filtering, and joint federated-centralized training—0
Proposal-based Few-shot Sound Event Detection for Speech and Environmental Sounds with Perceivers—0
Prototype-based Personalized Pruning—0
Prototypical Metric Transfer Learning for Continuous Speech Keyword Spotting With Limited Training Data—0
PSCNN: A 885.86 TOPS/W Programmable SRAM-based Computing-In-Memory Processor for Keyword Spotting—0
QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using MLPMixer—0
Query-by-Example Keyword Spotting system using Multi-head Attention and Softtriple Loss—0
Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning—0
Query-by-example on-device keyword spotting—0
ReckOn: A 28nm Sub-mm2 Task-Agnostic Spiking Recurrent Neural Network Processor Enabling On-Chip Learning over Second-Long Timescales—0
Recycle Your Wav2Vec2 Codebook: A Speech Perceiver for Keyword Spotting—0
Reformulating Information Retrieval from Speech and Text as a Detection Problem—0
Relational Proxy Loss for Audio-Text based Keyword Spotting—0
RepCNN: Micro-sized, Mighty Models for Wakeword Detection—0
Resource-Efficient Neural Architect—0
RNNAccel: A Fusion Recurrent Neural Network Accelerator for Edge Intelligence—0
Robust Spoken Term Detection Automatically Adjusted for a Given Threshold—0
Scalable Weight Reparametrization for Efficient Transfer Learning—0
Self-supervised speech representation learning for keyword-spotting with light-weight transformers—0
Sequence Discriminative Training for Deep Learning based Acoustic Keyword Spotting—0
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting—0
Small-footprint Keyword Spotting Using Deep Neural Network and Connectionist Temporal Classifier—0
Small-footprint Keyword Spotting with Graph Convolutional Network—0
Small-Footprint Open-Vocabulary Keyword Spotting with Quantized LSTM Networks—0
Small-footprint slimmable networks for keyword spotting—0
Speech and language technologies for the automatic monitoring and training of cognitive functions—0
Speech Augmentation Based Unsupervised Learning for Keyword Spotting—0
Speech Enhancement for Wake-Up-Word detection in Voice Assistants—0
Speech-MLP: a simple MLP architecture for speech processing—0
Speech Privacy Leakage from Shared Gradients in Distributed Learning—0
Speech Recognition: Keyword Spotting Through Image Recognition—0
Speech Unlearning—0
SpeechYOLO: Detection and Localization of Speech Objects—0
Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks—0
Split Federated Learning on Micro-controllers: A Keyword Spotting Showcase—0
Show:102550
← PrevPage 7 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NNI non-filtered(for the development set)Cnxe6.09—Unverified
2NNI Choi(for the development set)Cnxe5.89—Unverified
3NTU rnn (eval)Cnxe2.01—Unverified
4NTU dtw (eval)Cnxe2.01—Unverified
5NTU dtw (dev)Cnxe2.01—Unverified
6NTU rnn (dev)Cnxe2.01—Unverified
7ELiRF SDTW (eval)Cnxe1.19—Unverified
8ELiRF SDTW-avg (eval)Cnxe1.07—Unverified
9ELiRF SDTW (dev)Cnxe1.07—Unverified
10CUNY [Subseq+MFCC] (eval)Cnxe1.07—Unverified
#ModelMetricClaimedVerifiedStatus
1WaveFormerGoogle Speech Commands V2 1298.8—Unverified
2QNNGoogle Speech Commands V2 3598.6—Unverified
3TripletLoss-res15Google Speech Commands V1 1298.56—Unverified
4M2DGoogle Speech Commands V2 3598.5—Unverified
5EAT-SGoogle Speech Commands V2 3598.15—Unverified
6Audio Spectrogram TransformerGoogle Speech Commands V2 3598.11—Unverified
7EdgeCRNN 2.0×Google Speech Commands V2 1298.05—Unverified
8BC-ResNet-8Google Speech Commands V1 1298—Unverified
9HTS-ATGoogle Speech Commands V2 3598—Unverified
10Wav2KWSGoogle Speech Commands V1 1297.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Stacked 1D CNNError Rate1.99—Unverified
2End-to-end DNN-HMMError Rate1.7—Unverified
3HEiMDaLError Rate0.45—Unverified
#ModelMetricClaimedVerifiedStatus
1Res26Accuracy95.88—Unverified
2EfficientNet-A0 + SA + TLAccuracy95.83—Unverified
#ModelMetricClaimedVerifiedStatus
1QuaternionNeuralNetworkAccuracy (10-fold)98.53—Unverified
2SSAMBAAccuracy (10-fold)97.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TensorFlow's model version 2TFMA89.7—Unverified
2TensorFlow's model version 1TFMA85.4—Unverified
#ModelMetricClaimedVerifiedStatus
12D-ConvNetAccuracy (%)95.4—Unverified
21D-ConvNetAccuracy (%)93.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Quaternion Neural NetworksAccuracy(10-fold)98.53—Unverified
#ModelMetricClaimedVerifiedStatus
1MicroNet-KWS-LAccuracy95.3—Unverified