| S^3: Synonymous Semantic Space for Improving Zero-Shot Generalization of Vision-Language Models | Dec 6, 2024 | zero-shot-classificationZero-shot Generalization | —Unverified | 0 |
| Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail | Dec 5, 2024 | Stereo MatchingZero-shot Generalization | CodeCode Available | 3 |
| CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance | Dec 5, 2024 | Contrastive Learningcross-modal alignment | —Unverified | 0 |
| UTSD: Unified Time Series Diffusion Model | Dec 4, 2024 | Denoisingmodel | —Unverified | 0 |
| The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control | Dec 4, 2024 | Zero-shot Generalization | —Unverified | 0 |
| COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection | Nov 28, 2024 | object-detectionObject Detection | CodeCode Available | 1 |
| Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient | Nov 26, 2024 | GPUImage Generation | CodeCode Available | 2 |
| Visatronic: A Multimodal Decoder-Only Model for Speech Synthesis | Nov 26, 2024 | Decodermultimodal generation | —Unverified | 0 |
| vesselFM: A Foundation Model for Universal 3D Blood Vessel Segmentation | Nov 26, 2024 | Image SegmentationMedical Image Analysis | CodeCode Available | 2 |
| Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models | Nov 25, 2024 | Domain GeneralizationPrompt Learning | —Unverified | 0 |