| Active Data Curation Effectively Distills Large-Scale Multimodal Models | Nov 27, 2024 | DecoderImage Captioning | —Unverified | 0 |
| ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering? | Nov 27, 2024 | Question AnsweringVisual Question Answering | —Unverified | 0 |
| Task Progressive Curriculum Learning for Robust Visual Question Answering | Nov 26, 2024 | Data AugmentationEnsemble Learning | —Unverified | 0 |
| Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey | Nov 26, 2024 | Natural Language UnderstandingQuestion Answering | —Unverified | 0 |
| Efficient Multi-modal Large Language Models via Visual Token Grouping | Nov 26, 2024 | Image CaptioningQuestion Answering | —Unverified | 0 |
| GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis | Nov 25, 2024 | Medical Visual Question AnsweringMultiple-choice | —Unverified | 0 |
| Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering | Nov 24, 2024 | Question AnsweringRelational Reasoning | —Unverified | 0 |
| ReWind: Understanding Long Videos with Instructed Learnable Memory | Nov 23, 2024 | Large Language ModelQuestion Answering | —Unverified | 0 |
| freePruner: A Training-free Approach for Large Multimodal Model Acceleration | Nov 23, 2024 | QuantizationQuestion Answering | —Unverified | 0 |
| FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity | Nov 23, 2024 | AttributeCross-Modal Retrieval | —Unverified | 0 |
| Visual Contexts Clarify Ambiguous Expressions: A Benchmark Dataset | Nov 21, 2024 | Question AnsweringVisual Grounding | CodeCode Available | 0 |
| FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression | Nov 21, 2024 | Visual Question Answering | —Unverified | 0 |
| Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance | Nov 21, 2024 | Visual Question Answering | —Unverified | 0 |
| LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement | Nov 20, 2024 | Autonomous DrivingComputational Efficiency | —Unverified | 0 |
| Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training | Nov 20, 2024 | Contrastive Learningimage-classification | —Unverified | 0 |
| Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model | Nov 19, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 |
| CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs | Nov 19, 2024 | HallucinationLanguage Modeling | —Unverified | 0 |
| Value-Spectrum: Quantifying Preferences of Vision-Language Models via Value Decomposition in Social Media Contexts | Nov 18, 2024 | BenchmarkingMultimodal Large Language Model | CodeCode Available | 0 |
| Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering | Nov 17, 2024 | HallucinationIn-Context Learning | CodeCode Available | 0 |
| Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry | Nov 17, 2024 | Question AnsweringScene Understanding | —Unverified | 0 |
| A Comprehensive Survey on Visual Question Answering Datasets and Algorithms | Nov 17, 2024 | DiagnosticMiscellaneous | —Unverified | 0 |
| Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning | Nov 17, 2024 | Image CaptioningLanguage Modeling | CodeCode Available | 0 |
| Large Vision-Language Models for Remote Sensing Visual Question Answering | Nov 16, 2024 | Language ModelingLanguage Modelling | —Unverified | 0 |
| Visual question answering based evaluation metrics for text-to-image generation | Nov 15, 2024 | Image GenerationImage Manipulation | —Unverified | 0 |
| AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference | Nov 15, 2024 | QuantizationQuestion Answering | —Unverified | 0 |