| Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images | Aug 31, 2023 | 3D Shape GenerationContrastive Learning | —Unverified | 0 |
| DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment | Aug 22, 2023 | AttributeConstituency Parsing | —Unverified | 0 |
| Language-Guided Diffusion Model for Visual Grounding | Aug 18, 2023 | cross-modal alignmentDenoising | CodeCode Available | 0 |
| Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation | Aug 2, 2023 | cross-modal alignmentDenoising | CodeCode Available | 0 |
| WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation | Jun 19, 2023 | cross-modal alignmentImage Segmentation | —Unverified | 0 |
| Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models | Jun 15, 2023 | cross-modal alignmentDomain Generalization | —Unverified | 0 |
| Improving speech translation by fusing speech and text | May 23, 2023 | cross-modal alignmentMachine Translation | —Unverified | 0 |
| Speech-Text Dialog Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment | May 19, 2023 | cross-modal alignmentEmotion Recognition in Conversation | —Unverified | 0 |
| Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training | May 13, 2023 | cross-modal alignment | —Unverified | 0 |
| AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment | May 8, 2023 | cross-modal alignmentRhythm | —Unverified | 0 |
| CoVLR: Coordinating Cross-Modal Consistency and Intra-Modal Structure for Vision-Language Retrieval | Apr 15, 2023 | cross-modal alignmentCross-Modal Retrieval | —Unverified | 0 |
| SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger | Mar 30, 2023 | cross-modal alignmentzero-shot-classification | —Unverified | 0 |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models | Mar 28, 2023 | Action ClassificationAction Recognition | CodeCode Available | 0 |
| LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion | Mar 7, 2023 | 3D Object Detectioncross-modal alignment | CodeCode Available | 0 |
| TOT: Topology-Aware Optimal Transport For Multimodal Hate Detection | Feb 27, 2023 | cross-modal alignment | —Unverified | 0 |
| End-to-end Semantic Object Detection with Cross-Modal Alignment | Feb 10, 2023 | Contrastive Learningcross-modal alignment | —Unverified | 0 |
| Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? | Feb 1, 2023 | cross-modal alignmentLanguage Acquisition | —Unverified | 0 |
| Improving Cross-modal Alignment for Text-Guided Image Inpainting | Jan 26, 2023 | cross-modal alignmentImage Inpainting | —Unverified | 0 |
| Linguistic Query-Guided Mask Generation for Referring Image Segmentation | Jan 16, 2023 | Contrastive Learningcross-modal alignment | —Unverified | 0 |
| HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training | Dec 30, 2022 | cross-modal alignmentTGIF-Action | —Unverified | 0 |
| SimVTP: Simple Video Text Pre-training with Masked Autoencoders | Dec 7, 2022 | Contrastive Learningcross-modal alignment | CodeCode Available | 0 |
| Asymmetric Cross-Scale Alignment for Text-Based Person Search | Nov 26, 2022 | cross-modal alignmentPerson Search | CodeCode Available | 0 |
| How do Cross-View and Cross-Modal Alignment Affect Representations in Contrastive Learning? | Nov 23, 2022 | Contrastive Learningcross-modal alignment | —Unverified | 0 |
| SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training | Nov 21, 2022 | cross-modal alignmentGPU | —Unverified | 0 |
| Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision | Oct 24, 2022 | cross-modal alignmentCross-Modal Retrieval | —Unverified | 0 |