SOTAVerified

Automatic Speech Recognition (ASR)

Automatic Speech Recognition (ASR) involves converting spoken language into written text. It is designed to transcribe spoken words into text in real-time, allowing people to communicate with computers, mobile devices, and other technology using their voice. The goal of Automatic Speech Recognition is to accurately transcribe speech, taking into account variations in accent, pronunciation, and speaking style, as well as background noise and other factors that can affect speech quality.

Papers

Showing 9511000 of 3012 papers

TitleStatusHype
Don't Be So Sure! Boosting ASR Decoding via Confidence Relaxation0
Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters0
Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?0
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio0
DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised Learning and Its Application to Children's ASR0
Driving ROVER with Segment-based ASR Quality Estimation0
An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions0
Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition0
On the Effectiveness of Pinyin-Character Dual-Decoding for End-to-End Mandarin Chinese ASR0
Can You Hear It? Backdoor Attacks via Ultrasonic Triggers0
Dual Language Models for Code Switched Speech Recognition0
Can Whisper perform speech-based in-context learning?0
Dual Script E2E framework for Multilingual and Code-Switching ASR0
DUAL: Textless Spoken Question Answering with Speech Discrete Unit Adaptive Learning0
Are disentangled representations all you need to build speaker anonymization systems?0
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion0
DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input0
DuDe: Dual-Decoder Multilingual ASR for Indian Languages using Common Label Set0
Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition0
DuTongChuan: Context-aware Translation Model for Simultaneous Interpreting0
Dynamic Acoustic Unit Augmentation With BPE-Dropout for Low-Resource End-to-End Speech Recognition0
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model0
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization0
Dynamic Data Pruning for Automatic Speech Recognition0
Dynamic Latency for CTC-Based Streaming Automatic Speech Recognition With Emformer0
Dynamic Masking for Improved Stability in Spoken Language Translation0
Can We Trust Explainable AI Methods on ASR? An Evaluation on Phoneme Recognition0
Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification0
Early Stage LM Integration Using Local and Global Log-Linear Combination0
A Recorded Debating Dataset0
Can We Train a Language Model Inside an End-to-End ASR Model? - Investigating Effective Implicit Language Modeling0
E-Branchformer: Branchformer with Enhanced merging for speech recognition0
Echo State Speech Recognition0
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks0
A Vietnamese Dialog Act Corpus Based on ISO 24617-2 standard0
EdgeCRNN: an edgecomputing oriented model of acoustic feature enhancement for keyword spotting0
探討聲學模型的合併技術與半監督鑑別式訓練於會議語音辨識之研究 (Investigating acoustic model combination and semi-supervised discriminative training for meeting speech recognition) [In Chinese]0
會議語音辨識使用語者資訊之語言模型調適技術 (On the Use of Speaker-Aware Language Model Adaptation Techniques for Meeting Speech Recognition ) [In Chinese]0
Can Visual Context Improve Automatic Speech Recognition for an Embodied Agent?0
EEG based Continuous Speech Recognition using Transformers0
Adversarial Speaker Adaptation0
Effectively pretraining a speech translation decoder with Machine Translation data0
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition0
Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition0
Cantonese Automatic Speech Recognition Using Transfer Learning from Mandarin0
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning0
Effect of noise suppression losses on speech distortion and ASR performance0
Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models0
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech0
A Wav2vec2-Based Experimental Study on Self-Supervised Learning Methods to Improve Child Speech Recognition0
Show:102550
← PrevPage 20 of 61Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TM-CTCTest WER10.1Unverified
2TM-seq2seqTest WER9.7Unverified
3CTC/attentionTest WER8.2Unverified
4LF-MMI TDNNTest WER6.7Unverified
5Whisper-LLaMATest WER6.6Unverified
6End2end ConformerTest WER3.9Unverified
7End2end ConformerTest WER3.7Unverified
8MoCo + wav2vec (w/o extLM)Test WER2.7Unverified
9CTC/AttentionTest WER1.5Unverified
10WhisperTest WER1.3Unverified
#ModelMetricClaimedVerifiedStatus
1SpatialNetCER14.5Unverified
2CleanMel-L-maskCER14.4Unverified
#ModelMetricClaimedVerifiedStatus
1ConformerTest WER15.32Unverified
2Whisper-largev3-finetunedTest WER10.82Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)1.89Unverified
#ModelMetricClaimedVerifiedStatus
1DistillAVWER1.4Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)4.28Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)8.04Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)3.36Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer Transducer (German)WER (%)8.98Unverified