| Probing Visual Language Priors in VLMs | Dec 31, 2024 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models | Dec 31, 2024 | Multiple-choiceQuestion Answering | CodeCode Available | 0 |
| Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering | Dec 30, 2024 | Image CaptioningObject Recognition | —Unverified | 0 |
| UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models | Dec 30, 2024 | Question AnsweringScene Classification | CodeCode Available | 0 |
| HALLUCINOGEN: A Benchmark for Evaluating Object Hallucination in Large Visual-Language Models | Dec 29, 2024 | HallucinationObject | CodeCode Available | 0 |
| ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers | Dec 27, 2024 | Image CaptioningQuestion Answering | —Unverified | 0 |
| LININ: Logic Integrated Neural Inference Network for Explanatory Visual Question Answering | Dec 24, 2024 | Explanatory Visual Question AnsweringMultimodal Reasoning | CodeCode Available | 0 |
| Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering | Dec 24, 2024 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization | Dec 24, 2024 | In-Context LearningQuestion Answering | —Unverified | 0 |
| Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy | Dec 23, 2024 | Image CaptioningQuestion Answering | —Unverified | 0 |
| Multimodal Preference Data Synthetic Alignment with Reward Model | Dec 23, 2024 | 2kCaption Generation | CodeCode Available | 0 |
| Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective | Dec 23, 2024 | Question AnsweringVisual Question Answering | CodeCode Available | 0 |
| FFA Sora, video generation as fundus fluorescein angiography simulator | Dec 23, 2024 | Privacy PreservingQuestion Answering | —Unverified | 0 |
| Prompting Large Language Models with Rationale Heuristics for Knowledge-based Visual Question Answering | Dec 22, 2024 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization | Dec 21, 2024 | Image CaptioningMultimodal Reasoning | CodeCode Available | 0 |
| NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional Generalization | Dec 20, 2024 | Compositional Generalization (AVG)Novel Concepts | CodeCode Available | 0 |
| FedPIA -- Permuting and Integrating Adapters leveraging Wasserstein Barycenters for Finetuning Foundation Models in Multi-Modal Federated Learning | Dec 19, 2024 | Federated Learningparameter-efficient fine-tuning | —Unverified | 0 |
| Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models | Dec 19, 2024 | Autonomous DrivingImage Captioning | CodeCode Available | 0 |
| Consistency of Compositional Generalization across Multiple Levels | Dec 18, 2024 | Meta-LearningQuestion Answering | CodeCode Available | 0 |
| A Concept-Centric Approach to Multi-Modality Learning | Dec 18, 2024 | Image-text matchingQuestion Answering | —Unverified | 0 |
| Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues | Dec 17, 2024 | Language ModelingLanguage Modelling | CodeCode Available | 0 |
| CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology | Dec 16, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 |
| LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering | Dec 16, 2024 | In-Context LearningInstruction Following | CodeCode Available | 0 |
| Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track | Dec 15, 2024 | Image CaptioningMedical Question Answering | —Unverified | 0 |
| Damage Assessment after Natural Disasters with UAVs: Semantic Feature Extraction using Deep Learning | Dec 14, 2024 | Decision MakingQuestion Answering | —Unverified | 0 |
| Patch-level Sounding Object Tracking for Audio-Visual Question Answering | Dec 14, 2024 | Audio-visual Question AnsweringObject Tracking | —Unverified | 0 |
| VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation | Dec 13, 2024 | Instruction FollowingQuestion Answering | —Unverified | 0 |
| ViUniT: Visual Unit Tests for More Robust Visual Programming | Dec 12, 2024 | Image GenerationImage-text matching | —Unverified | 0 |
| Discrete Subgraph Sampling for Interpretable Graph based Visual Question Answering | Dec 11, 2024 | Explainable artificial intelligenceExplainable Artificial Intelligence (XAI) | CodeCode Available | 0 |
| Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions | Dec 11, 2024 | BenchmarkingQuestion Answering | CodeCode Available | 0 |
| Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses | Dec 11, 2024 | Image-text RetrievalQuestion Answering | —Unverified | 0 |
| How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey | Dec 11, 2024 | Image CaptioningQuestion Answering | —Unverified | 0 |
| A Multimodal Social Agent | Dec 11, 2024 | Common Sense ReasoningDecision Making | —Unverified | 0 |
| Can We Generate Visual Programs Without Prompting LLMs? | Dec 11, 2024 | Data AugmentationQuestion Answering | —Unverified | 0 |
| MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models | Dec 10, 2024 | Multiple-choiceQuestion Answering | CodeCode Available | 0 |
| ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance | Dec 9, 2024 | Image GenerationLanguage Modeling | —Unverified | 0 |
| Ranked from Within: Ranking Large Multimodal Models for Visual Question Answering Without Labels | Dec 9, 2024 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering | Dec 9, 2024 | Knowledge DistillationQuestion Answering | CodeCode Available | 0 |
| Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora | Dec 6, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 |
| Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling | Dec 6, 2024 | document understandingHallucination | —Unverified | 0 |
| EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation | Dec 6, 2024 | MMEQuestion Answering | —Unverified | 0 |
| T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts | Dec 5, 2024 | BenchmarkingImage Generation | —Unverified | 0 |
| Copy-Move Forgery Detection and Question Answering for Remote Sensing Image | Dec 3, 2024 | Question AnsweringVisual Question Answering | CodeCode Available | 0 |
| Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey | Dec 3, 2024 | Cross-Modal RetrievalNatural Language Understanding | —Unverified | 0 |
| CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs | Dec 3, 2024 | Image CaptioningQuantization | —Unverified | 0 |
| Understanding the World's Museums through Vision-Language Reasoning | Dec 2, 2024 | BenchmarkingQuestion Answering | CodeCode Available | 0 |
| DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness | Nov 29, 2024 | Optical Character Recognition (OCR)Question Answering | CodeCode Available | 0 |
| SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks | Nov 29, 2024 | Question AnsweringVisual Question Answering | CodeCode Available | 0 |
| Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs | Nov 28, 2024 | AttributeHallucination | —Unverified | 0 |
| Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers | Nov 28, 2024 | Image Captioningimage-classification | —Unverified | 0 |