| Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval | Dec 19, 2023 | cross-modal alignmentMoment Retrieval | CodeCode Available | 1 |
| Mask Grounding for Referring Image Segmentation | Dec 19, 2023 | cross-modal alignmentImage Segmentation | CodeCode Available | 1 |
| M^2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge Base | Dec 16, 2023 | cross-modal alignmentKnowledge Graphs | CodeCode Available | 0 |
| Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning | Dec 14, 2023 | cross-modal alignmentDecoder | —Unverified | 0 |
| ViLA: Efficient Video-Language Alignment for Video Question Answering | Dec 13, 2023 | cross-modal alignmentLanguage Modeling | CodeCode Available | 1 |
| OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection | Dec 12, 2023 | cross-modal alignmentobject-detection | —Unverified | 0 |
| Navigating Open Set Scenarios for Skeleton-based Action Recognition | Dec 11, 2023 | Action RecognitionActivity Recognition | CodeCode Available | 1 |
| Progressive Multi-Modality Learning for Inverse Protein Folding | Dec 11, 2023 | cross-modal alignmentData Augmentation | CodeCode Available | 1 |
| PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features | Dec 5, 2023 | cross-modal alignmentDecoder | —Unverified | 0 |
| DAP: Domain-aware Prompt Learning for Vision-and-Language Navigation | Nov 29, 2023 | cross-modal alignmentNavigate | —Unverified | 0 |