SOTAVerified

Automatic Speech Recognition (ASR)

Automatic Speech Recognition (ASR) involves converting spoken language into written text. It is designed to transcribe spoken words into text in real-time, allowing people to communicate with computers, mobile devices, and other technology using their voice. The goal of Automatic Speech Recognition is to accurately transcribe speech, taking into account variations in accent, pronunciation, and speaking style, as well as background noise and other factors that can affect speech quality.

Papers

Showing 9761000 of 3012 papers

TitleStatusHype
Dynamic Masking for Improved Stability in Spoken Language Translation0
Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition0
Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification0
Early Stage LM Integration Using Local and Global Log-Linear Combination0
A Recorded Debating Dataset0
Can We Train a Language Model Inside an End-to-End ASR Model? - Investigating Effective Implicit Language Modeling0
E-Branchformer: Branchformer with Enhanced merging for speech recognition0
Echo State Speech Recognition0
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks0
A Vietnamese Dialog Act Corpus Based on ISO 24617-2 standard0
EdgeCRNN: an edgecomputing oriented model of acoustic feature enhancement for keyword spotting0
探討聲學模型的合併技術與半監督鑑別式訓練於會議語音辨識之研究 (Investigating acoustic model combination and semi-supervised discriminative training for meeting speech recognition) [In Chinese]0
會議語音辨識使用語者資訊之語言模型調適技術 (On the Use of Speaker-Aware Language Model Adaptation Techniques for Meeting Speech Recognition ) [In Chinese]0
Can Visual Context Improve Automatic Speech Recognition for an Embodied Agent?0
EEG based Continuous Speech Recognition using Transformers0
Adversarial Speaker Adaptation0
Effectively pretraining a speech translation decoder with Machine Translation data0
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition0
Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition0
Cantonese Automatic Speech Recognition Using Transfer Learning from Mandarin0
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning0
Effect of noise suppression losses on speech distortion and ASR performance0
Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models0
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech0
A Wav2vec2-Based Experimental Study on Self-Supervised Learning Methods to Improve Child Speech Recognition0
Show:102550
← PrevPage 40 of 121Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TM-CTCTest WER10.1Unverified
2TM-seq2seqTest WER9.7Unverified
3CTC/attentionTest WER8.2Unverified
4LF-MMI TDNNTest WER6.7Unverified
5Whisper-LLaMATest WER6.6Unverified
6End2end ConformerTest WER3.9Unverified
7End2end ConformerTest WER3.7Unverified
8MoCo + wav2vec (w/o extLM)Test WER2.7Unverified
9CTC/AttentionTest WER1.5Unverified
10WhisperTest WER1.3Unverified
#ModelMetricClaimedVerifiedStatus
1SpatialNetCER14.5Unverified
2CleanMel-L-maskCER14.4Unverified
#ModelMetricClaimedVerifiedStatus
1ConformerTest WER15.32Unverified
2Whisper-largev3-finetunedTest WER10.82Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)1.89Unverified
#ModelMetricClaimedVerifiedStatus
1DistillAVWER1.4Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)4.28Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)8.04Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)3.36Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer Transducer (German)WER (%)8.98Unverified