| Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation | Aug 2, 2023 | cross-modal alignmentDenoising | CodeCode Available | 0 |
| Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval | Jul 18, 2023 | cross-modal alignmentData Augmentation | CodeCode Available | 1 |
| WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation | Jun 19, 2023 | cross-modal alignmentImage Segmentation | —Unverified | 0 |
| Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models | Jun 15, 2023 | cross-modal alignmentDomain Generalization | —Unverified | 0 |
| Global and Local Semantic Completion Learning for Vision-Language Pre-training | Jun 12, 2023 | cross-modal alignmentImage-text Retrieval | CodeCode Available | 1 |
| ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning | May 31, 2023 | cross-modal alignmentRepresentation Learning | CodeCode Available | 1 |
| SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation | May 26, 2023 | cross-modal alignmentObject | CodeCode Available | 1 |
| Improving speech translation by fusing speech and text | May 23, 2023 | cross-modal alignmentMachine Translation | —Unverified | 0 |
| Speech-Text Dialog Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment | May 19, 2023 | cross-modal alignmentEmotion Recognition in Conversation | —Unverified | 0 |
| Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training | May 13, 2023 | cross-modal alignment | —Unverified | 0 |