SOTAVerified

MME

MME is a comprehensive evaluation benchmark for multimodal large language models. It measures both perception and cognition abilities on a total of 14 subtasks, including existence, count, position, color, poster, celebrity, scene, landmark, artwork, OCR, commonsense reasoning, numerical calculation, text translation, and code reasoning.

Papers

Showing 51–75 of 95 papers

TitleStatusHype
A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise—0
AIDE: Agentically Improve Visual Language Model with Domain Experts—0
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes—0
Apollo: An Exploration of Video Understanding in Large Multimodal Models—0
Benchmarking and In-depth Performance Study of Large Language Models on Habana Gaudi Processors—0
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination—0
Deep Learning for Hybrid 5G Services in Mobile Edge Computing Systems: Learn from a Digital Twin—0
Domain Adaptation via Minimax Entropy for Real/Bogus Classification of Astronomical Alerts—0
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models—0
DrVideo: Document Retrieval Based Long Video Understanding—0
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding—0
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation—0
The economic value of empowering older patients transitioning from hospital to home: Evidence from the 'Your Care Needs You' intervention—0
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy—0
Enhancing the Spatial Awareness Capability of Multi-Modal Large Language Model—0
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models—0
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models—0
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering—0
GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors—0
Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment—0
Improving LLM Video Understanding with 16 Frames Per Second—0
Language-Vision Planner and Executor for Text-to-Visual Reasoning—0
Learning Multilingual Meta-Embeddings for Code-Switching Named Entity Recognition—0
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding—0
Machine Learning Methods for Inferring the Number of UAV Emitters via Massive MIMO Receive Array—0
Show:102550
← PrevPage 3 of 4Next →

No leaderboard results yet.