| DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module | Aug 31, 2024 | Audio-Visual Speech Recognitionspeech-recognition | —Unverified | 0 |
| The NPU-ASLP System Description for Visual Speech Recognition in CNVSRC 2024 | Aug 5, 2024 | Decoderspeech-recognition | CodeCode Available | 0 |
| SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data | Aug 1, 2024 | Audio-Visual Speech RecognitionAutomatic Speech Recognition | CodeCode Available | 0 |
| Tailored Design of Audio-Visual Speech Recognition Models using Branchformers | Jul 9, 2024 | Audio-Visual Speech Recognitionspeech-recognition | CodeCode Available | 1 |
| Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition | Jul 4, 2024 | Audio-Visual Speech Recognitionspeech-recognition | CodeCode Available | 1 |
| MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization | Jun 25, 2024 | Audio-Visual Speech Recognitionspeech-recognition | —Unverified | 0 |
| SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization | Jun 18, 2024 | Landmark-based LipreadingLipreading | CodeCode Available | 2 |
| CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge | Jun 14, 2024 | speech-recognitionSpeech Recognition | —Unverified | 0 |
| Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation | Jun 14, 2024 | Audio-Visual Speech RecognitionAutomatic Speech Recognition (ASR) | CodeCode Available | 3 |
| Watch Your Mouth: Silent Speech Recognition with Depth Sensing | May 11, 2024 | Deep LearningLipreading | CodeCode Available | 1 |