SOTAVerified

Image Comprehension

Papers

Showing 4149 of 49 papers

TitleStatusHype
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition0
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine0
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image InputsCode0
GeoLocator: a location-integrated large multimodal model for inferring geo-privacy0
What Large Language Models Bring to Text-rich VQA?0
On the Performance of Multimodal Language Models0
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and CompositionCode0
Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens0
An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension0
Show:102550
← PrevPage 5 of 5Next →

No leaderboard results yet.