| One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning | Aug 6, 2024 | AllImage Captioning | —Unverified | 0 |
| Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins | Mar 26, 2025 | Large Language ModelReasoning Segmentation | —Unverified | 0 |
| Transferring Foundation Models for Generalizable Robotic Manipulation | Jun 9, 2023 | Imitation LearningObject | —Unverified | 0 |
| VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation | Mar 18, 2025 | Reasoning SegmentationVideo Editing | —Unverified | 0 |
| Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level | Nov 15, 2024 | Benchmarkingcounterfactual | —Unverified | 0 |
| MediSee: Reasoning-based Pixel-level Perception in Medical Images | Apr 15, 2025 | Logical ReasoningReasoning Segmentation | —Unverified | 0 |
| LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation | Apr 15, 2025 | Image CaptioningQuestion Answering | —Unverified | 0 |
| PixelThink: Towards Efficient Chain-of-Pixel Reasoning | May 29, 2025 | Reasoning Segmentationreinforcement-learning | —Unverified | 0 |
| POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation | Jan 1, 2025 | HallucinationReasoning Segmentation | —Unverified | 0 |
| PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation | Dec 19, 2024 | Reasoning Segmentation | —Unverified | 0 |
| PRS-Med: Position Reasoning Segmentation with Vision-Language Model in Medical Imaging | May 17, 2025 | Image SegmentationLanguage Modeling | —Unverified | 0 |
| LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery | May 5, 2025 | Reasoning SegmentationSegmentation | —Unverified | 0 |
| Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models | May 29, 2024 | 3D Instance Segmentation3D Semantic Segmentation | —Unverified | 0 |
| Reasoning Segmentation for Images and Videos: A Survey | May 24, 2025 | Reasoning SegmentationSurvey | —Unverified | 0 |
| RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought | Jun 4, 2025 | Multimodal ReasoningReasoning Segmentation | —Unverified | 0 |
| SegLLM: Multi-round Reasoning Segmentation | Oct 24, 2024 | Reasoning SegmentationReferring Expression | —Unverified | 0 |
| HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation | Jul 17, 2025 | Reasoning SegmentationWorld Knowledge | —Unverified | 0 |
| FoodLMM: A Versatile Food Assistant using Large Multi-modal Model | Dec 22, 2023 | Food RecognitionMulti-Task Learning | —Unverified | 0 |
| Beyond Segmentation: Road Network Generation with Multi-Modal LLMs | Oct 15, 2023 | Autonomous NavigationLanguage Modeling | —Unverified | 0 |
| Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts | Mar 10, 2025 | Reasoning SegmentationSegmentation | —Unverified | 0 |
| Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations | Jun 9, 2025 | Large Language ModelMultimodal Reasoning | —Unverified | 0 |
| Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA | Mar 13, 2025 | Dataset GenerationReasoning Segmentation | —Unverified | 0 |
| MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models | Jun 12, 2025 | Image SegmentationMedical Diagnosis | —Unverified | 0 |
| MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation | Mar 23, 2025 | Language ModelingLanguage Modelling | —Unverified | 0 |
| Pixel-Level Reasoning Segmentation via Multi-turn Conversations | Feb 13, 2025 | Reasoning SegmentationSegmentation | CodeCode Available | 0 |