SOTAVerified

Automatic Speech Recognition (ASR)

Automatic Speech Recognition (ASR) involves converting spoken language into written text. It is designed to transcribe spoken words into text in real-time, allowing people to communicate with computers, mobile devices, and other technology using their voice. The goal of Automatic Speech Recognition is to accurately transcribe speech, taking into account variations in accent, pronunciation, and speaking style, as well as background noise and other factors that can affect speech quality.

Papers

Showing 9511000 of 3012 papers

TitleStatusHype
Don't Be So Sure! Boosting ASR Decoding via Confidence Relaxation0
Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters0
Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?0
Do You Listen with One or Two Microphones? A Unified ASR Model for Single and Multi-Channel Audio0
DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised Learning and Its Application to Children's ASR0
Driving ROVER with Segment-based ASR Quality Estimation0
An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions0
Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition0
On the Effectiveness of Pinyin-Character Dual-Decoding for End-to-End Mandarin Chinese ASR0
Contrastive Semi-supervised Learning for ASR0
Dual Language Models for Code Switched Speech Recognition0
ATCSpeechNet: A multilingual end-to-end speech recognition framework for air traffic control systems0
Dual Script E2E framework for Multilingual and Code-Switching ASR0
DUAL: Textless Spoken Question Answering with Speech Discrete Unit Adaptive Learning0
Learning Video Representations using Contrastive Bidirectional Transformer0
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion0
DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input0
DuDe: Dual-Decoder Multilingual ASR for Indian Languages using Common Label Set0
ATCSpeech: a multilingual pilot-controller speech corpus from real Air Traffic Control environment0
DuTongChuan: Context-aware Translation Model for Simultaneous Interpreting0
Dynamic Acoustic Unit Augmentation With BPE-Dropout for Low-Resource End-to-End Speech Recognition0
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model0
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization0
Dynamic Data Pruning for Automatic Speech Recognition0
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR0
Dynamic Masking for Improved Stability in Spoken Language Translation0
Acoustic Word Disambiguation with Phonogical Features in Danish ASR0
Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification0
Continuous Speech Recognition using EEG and Video0
Continuous Pseudo-Labeling from the Start0
Continuously Learning New Words in Automatic Speech Recognition0
E-Branchformer: Branchformer with Enhanced merging for speech recognition0
ATC-ANNO: Semantic Annotation for Air Traffic Control with Assistive Auto-Annotation0
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks0
Algorithms For Automatic Accentuation And Transcription Of Russian Texts In Speech Recognition Systems0
EdgeCRNN: an edgecomputing oriented model of acoustic feature enhancement for keyword spotting0
探討聲學模型的合併技術與半監督鑑別式訓練於會議語音辨識之研究 (Investigating acoustic model combination and semi-supervised discriminative training for meeting speech recognition) [In Chinese]0
會議語音辨識使用語者資訊之語言模型調適技術 (On the Use of Speaker-Aware Language Model Adaptation Techniques for Meeting Speech Recognition ) [In Chinese]0
Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence0
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings0
An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation Learning0
Effectively pretraining a speech translation decoder with Machine Translation data0
Asynchronous Decentralized Distributed Training of Acoustic Models0
A Lexical-aware Non-autoregressive Transformer-based ASR Model0
Acoustic-to-articulatory Speech Inversion with Multi-task Learning0
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning0
Continual learning using lattice-free MMI for speech recognition0
Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models0
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory0
A Survey on Speech Large Language Models0
Show:102550
← PrevPage 20 of 61Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TM-CTCTest WER10.1Unverified
2TM-seq2seqTest WER9.7Unverified
3CTC/attentionTest WER8.2Unverified
4LF-MMI TDNNTest WER6.7Unverified
5Whisper-LLaMATest WER6.6Unverified
6End2end ConformerTest WER3.9Unverified
7End2end ConformerTest WER3.7Unverified
8MoCo + wav2vec (w/o extLM)Test WER2.7Unverified
9CTC/AttentionTest WER1.5Unverified
10WhisperTest WER1.3Unverified
#ModelMetricClaimedVerifiedStatus
1SpatialNetCER14.5Unverified
2CleanMel-L-maskCER14.4Unverified
#ModelMetricClaimedVerifiedStatus
1ConformerTest WER15.32Unverified
2Whisper-largev3-finetunedTest WER10.82Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)1.89Unverified
#ModelMetricClaimedVerifiedStatus
1DistillAVWER1.4Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)4.28Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)8.04Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer TransducerWER (%)3.36Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer Transducer (German)WER (%)8.98Unverified