SOTAVerified

Visual Speech Recognition

Papers

Showing 101–125 of 182 papers

TitleStatusHype
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision—0
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge—0
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video—0
Conformers are All You Need for Visual Speech Recognition—0
Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices—0
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition—0
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations—0
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset—0
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration—0
ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement—0
Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning—0
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning—0
Streaming Audio-Visual Speech Recognition with Alignment Regularization—0
Visual Speech Recognition in a Driver Assistance System—0
Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands—0
Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models—0
RUSAVIC Corpus: Russian Audio-Visual Speech in Cars—0
Is Lip Region-of-Interest Sufficient for Lipreading?—0
Deep Learning for Visual Speech Analysis: A Survey—0
Learning Contextually Fused Audio-visual Representations for Audio-visual Speech Recognition—0
Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Multi-Person Video—0
Recent Progress in the CUHK Dysarthric Speech Recognition System—0
Leveraging Uni-Modal Self-Supervised Learning for Multimodal Audio-visual Speech Recognition—0
Advances and Challenges in Deep Lip Reading—0
Sub-word Level Lip Reading With Visual Attention—0
Show:102550
← PrevPage 5 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)30.7—Unverified
2CTC/AttentionWord Error Rate (WER)19.1—Unverified
#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)22.6—Unverified