SOTAVerified

Visual Speech Recognition

Papers

Showing 51100 of 182 papers

TitleStatusHype
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition0
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023Code1
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationCode0
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition0
LiteVSR: Efficient Visual Speech Recognition by Learning from Speech Representations of Unlabeled Data0
The GUA-Speech System Description for CNVSRC Challenge 20230
Do VSR Models Generalize Beyond LRS3?Code1
Analysis of Visual Features for Continuous Lipreading in Spanish0
Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish0
LIP-RTVE: An Audiovisual Database for Continuous Spanish in the WildCode0
End-to-End Lip Reading in Romanian with Cross-Lingual Domain Adaptation and Lateral Inhibition0
AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition0
The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction0
Visual Speech Recognition for Languages with Limited Labeled Data using Automatic Labels from WhisperCode1
Another Point of View on Visual Speech Recognition0
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model0
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion EncoderCode1
Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping0
SparseVSR: Lightweight and Noise Robust Visual Speech Recognition0
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech RecognitionCode1
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionCode1
Automated Speaker Independent Visual Speech Recognition: A Comprehensive Survey0
OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentCode1
MAVD: The First Open Large-Scale Mandarin Audio-Visual Dataset with Depth InformationCode1
Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning0
Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task GeneralizationCode1
Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech RecognitionCode1
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition0
Deep Learning-based Spatio Temporal Facial Feature Visual Speech Recognition0
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision0
Auto-AVSR: Audio-Visual Speech Recognition with Automatic LabelsCode2
Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringCode1
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge0
MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionCode1
MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text TranslationCode2
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video0
Conformers are All You Need for Visual Speech Recognition0
Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices0
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition0
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations0
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset0
OLKAVS: An Open Large-Scale Korean Audio-Visual Speech DatasetCode1
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration0
ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement0
Jointly Learning Visual and Auditory Speech Representations from Raw DataCode1
Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning0
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning0
Streaming Audio-Visual Speech Recognition with Alignment Regularization0
Visual Speech Recognition in a Driver Assistance System0
Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech RecognitionCode1
Show:102550
← PrevPage 2 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)30.7Unverified
2CTC/AttentionWord Error Rate (WER)19.1Unverified
#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)22.6Unverified