| Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection | Jun 19, 2021 | Action Unit DetectionFacial Action Unit Detection | —Unverified | 0 |
| Image Change Captioning by Learning From an Auxiliary Task | Jun 19, 2021 | Image RetrievalMulti-Task Learning | —Unverified | 0 |
| Cascaded Prediction Network via Segment Tree for Temporal Video Grounding | Jun 19, 2021 | SentenceVideo Grounding | —Unverified | 0 |
| Sketch, Ground, and Refine: Top-Down Dense Video Captioning | Jun 19, 2021 | Dense Video CaptioningSentence | CodeCode Available | 0 |
| Structured Multi-Level Interaction Network for Video Moment Localization via Language Query | Jun 19, 2021 | Sentence | —Unverified | 0 |
| Multi-Stage Aggregated Transformer Network for Temporal Language Localization in Videos | Jun 19, 2021 | Sentence | —Unverified | 0 |
| Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading | Jun 19, 2021 | Lip ReadingSentence | —Unverified | 0 |
| Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event Captioning | Jun 19, 2021 | SentenceVideo Captioning | —Unverified | 0 |
| Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval | Jun 19, 2021 | Cross-Modal RetrievalGraph Matching | —Unverified | 0 |
| Transitional Adaptation of Pretrained Models for Visual Storytelling | Jun 19, 2021 | Image CaptioningLanguage Modelling | —Unverified | 0 |