SOTAVerified

Visual Speech Recognition

Papers

Showing 101150 of 182 papers

TitleStatusHype
SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision0
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge0
Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video0
Conformers are All You Need for Visual Speech Recognition0
Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices0
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition0
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations0
A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset0
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration0
ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement0
Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning0
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning0
Streaming Audio-Visual Speech Recognition with Alignment Regularization0
Visual Speech Recognition in a Driver Assistance System0
Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands0
Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models0
RUSAVIC Corpus: Russian Audio-Visual Speech in Cars0
Is Lip Region-of-Interest Sufficient for Lipreading?0
Deep Learning for Visual Speech Analysis: A Survey0
Learning Contextually Fused Audio-visual Representations for Audio-visual Speech Recognition0
Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Multi-Person Video0
Recent Progress in the CUHK Dysarthric Speech Recognition System0
Leveraging Uni-Modal Self-Supervised Learning for Multimodal Audio-visual Speech Recognition0
Advances and Challenges in Deep Lip Reading0
Sub-word Level Lip Reading With Visual Attention0
Perception Point: Identifying Critical Learning Periods in Speech for Bilingual Networks0
Audio-Visual Speech Recognition is Worth 32328 Voxels0
LRWR: Large-Scale Benchmark for Lip Reading in Russian language0
Large-vocabulary Audio-visual Speech Recognition in Noisy Environments0
Spatio-Temporal Attention Mechanism and Knowledge Distillation for Lip Reading0
Interactive decoding of words from visual speech recognition models0
Fusing information streams in end-to-end audio-visual speech recognition0
Part-based Lipreading for Audio-Visual Speech Recognition0
Lip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion0
"Notic My Speech" -- Blending Speech Patterns With Multimedia0
Audio-visual Recognition of Overlapped speech for the LRS2 dataset0
Detecting Adversarial Attacks On Audiovisual Speech Recognition0
Continuous Speech Recognition using EEG and Video0
ASR is all you need: cross-modal distillation for lip reading0
Recurrent Neural Network Transducer for Audio-Visual Speech RecognitionCode0
Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition0
MobiVSR: A Visual Speech Recognition Solution for Mobile Devices0
End-to-End Visual Speech Recognition for Small-Scale Datasets0
Harnessing GANs for Zero-shot Learning of New Classes in Visual Speech RecognitionCode0
Modality Attention for End-to-End Audio-visual Speech Recognition0
LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the WildCode0
3D Feature Pyramid Attention Module for Robust Visual Speech Recognition0
Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture0
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation0
LRS3-TED: a large-scale dataset for visual speech recognitionCode0
Show:102550
← PrevPage 3 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)30.7Unverified
2CTC/AttentionWord Error Rate (WER)19.1Unverified
#ModelMetricClaimedVerifiedStatus
1VTP with more dataWord Error Rate (WER)22.6Unverified