| SAMST: A Transformer framework based on SAM pseudo label filtering for remote sensing semi-supervised semantic segmentation | Jul 16, 2025 | Boundary DetectionPseudo Label | —Unverified | 0 |
| Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation | Jul 15, 2025 | 3D ReconstructionAutonomous Driving | —Unverified | 0 |
| PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment | Jul 12, 2025 | Large Language ModelPose Estimation | CodeCode Available | 0 |
| Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data | Jul 9, 2025 | Motion GenerationZero-shot Generalization | CodeCode Available | 0 |
| Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models | Jul 8, 2025 | Future predictionLarge Language Model | —Unverified | 0 |
| Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach | Jul 4, 2025 | AttributeContrastive Learning | —Unverified | 0 |
| DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment | Jul 3, 2025 | cross-modal alignmentInstruction Following | CodeCode Available | 2 |
| RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse Weather | Jul 2, 2025 | DenoisingDepth Estimation | —Unverified | 0 |
| WAFT: Warping-Alone Field Transforms for Optical Flow | Jun 26, 2025 | Optical Flow EstimationZero-shot Generalization | CodeCode Available | 2 |
| IRanker: Towards Ranking Foundation Model | Jun 25, 2025 | GSM8Kmodel | CodeCode Available | 1 |
| TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design | Jun 24, 2025 | Deep Reinforcement LearningZero-shot Generalization | CodeCode Available | 0 |
| VisLanding: Monocular 3D Perception for UAV Safe Landing via Depth-Normal Synergy | Jun 17, 2025 | Decision MakingSemantic Segmentation | —Unverified | 0 |
| LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction | Jun 16, 2025 | Instruction FollowingVision-Language-Action | —Unverified | 0 |
| Prohibited Items Segmentation via Occlusion-aware Bilayer Modeling | Jun 13, 2025 | DecoderImage Segmentation | CodeCode Available | 0 |
| DEAL: Disentangling Transformer Head Activations for LLM Steering | Jun 10, 2025 | Binary ClassificationZero-shot Generalization | —Unverified | 0 |
| Deep Equivariant Multi-Agent Control Barrier Functions | Jun 9, 2025 | Robot NavigationZero-shot Generalization | —Unverified | 0 |
| CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray | Jun 9, 2025 | ClassificationDiagnostic | —Unverified | 0 |
| ZeroVO: Visual Odometry with Minimal Assumptions | Jun 9, 2025 | Autonomous DrivingCamera Calibration | —Unverified | 0 |
| Latent Diffusion Model Based Denoising Receiver for 6G Semantic Communication: From Stochastic Differential Theory to Application | Jun 6, 2025 | DenoisingSemantic Communication | —Unverified | 0 |
| RecGPT: A Foundation Model for Sequential Recommendation | Jun 6, 2025 | Decodermodel | CodeCode Available | 2 |
| Towards Vision-Language-Garment Models For Web Knowledge Garment Understanding and Generation | Jun 5, 2025 | Zero-shot Generalization | —Unverified | 0 |
| Generating Synthetic Stereo Datasets using 3D Gaussian Splatting and Expert Knowledge Transfer | Jun 5, 2025 | 3DGSDataset Generation | —Unverified | 0 |
| OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis | Jun 4, 2025 | Action GenerationDecision Making | CodeCode Available | 1 |
| Language-Guided Multi-Agent Learning in Simulations: A Unified Framework and Evaluation | Jun 1, 2025 | Language ModelingLanguage Modelling | —Unverified | 0 |
| DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? | May 30, 2025 | DiagnosticMedical Image Analysis | CodeCode Available | 1 |