SOTAVerified

Lipreading

Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing.

Deep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks. It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc.

The primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.

Papers

Showing 1–10 of 103 papers

TitleStatusHype
Learning Speaker-Invariant Visual Features for Lipreading—0
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation—0
OXSeg: Multidimensional attention UNet-based lip segmentation using semi-supervised lip contours—0
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation—0
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation ModelsCode1
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data ConditionsCode0
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual InputsCode1
RAL:Redundancy-Aware Lipreading Model Based on Differential Learning with Symmetric Views—0
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token SynchronizationCode2
Watch Your Mouth: Silent Speech Recognition with Depth SensingCode1
Show:102550
← PrevPage 1 of 11Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1LIBSWord Error Rate (WER)65.29—Unverified
2TM-CTC + extLMWord Error Rate (WER)54.7—Unverified
3CTC + KD ASRWord Error Rate (WER)53.2—Unverified
4Conv-seq2seqWord Error Rate (WER)51.7—Unverified
5Hybrid CTC / AttentionWord Error Rate (WER)50—Unverified
6LF-MMI TDNNWord Error Rate (WER)48.86—Unverified
7TM-seq2seq + extLMWord Error Rate (WER)48.3—Unverified
8Multi-head Visual-Audio MemoryWord Error Rate (WER)44.5—Unverified
9MoCo + wav2vec (w/o extLM)Word Error Rate (WER)43.2—Unverified
10CTC/AttentionWord Error Rate (WER)32.9—Unverified