SOTAVerified

document understanding

Document understanding involves document classification, layout analysis, information extraction, and DocQA.

Papers

Showing 51100 of 309 papers

TitleStatusHype
Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsCode1
ARB: A Comprehensive Arabic Multimodal Reasoning BenchmarkCode1
MedICaT: A Dataset of Medical Images, Captions, and Textual ReferencesCode1
CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl DataCode1
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingCode1
DocLayLLM: An Efficient and Effective Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingCode1
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal LearningCode1
LineFormer: Rethinking Line Chart Data Extraction as Instance SegmentationCode1
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingCode1
DocFormerv2: Local Features for Document UnderstandingCode1
A Discrete Variational Recurrent Topic Model without the Reparametrization TrickCode1
Hierarchical Multimodal Pre-training for Visually Rich Webpage UnderstandingCode1
FRAG: Frame Selection Augmented Generation for Long Video and Long Document UnderstandingCode1
Doc2Graph: a Task Agnostic Document Understanding Framework based on Graph Neural NetworksCode1
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document UnderstandingCode1
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout TransformerCode1
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real WorldCode1
DocFormer: End-to-End Transformer for Document UnderstandingCode1
Multimodal Pre-training Based on Graph Attention Network for Document UnderstandingCode1
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning0
BERT-AL: BERT for Arbitrarily Long Document Understanding0
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review0
DeeperDive: The Unreasonable Effectiveness of Weak Supervision in Document Understanding A Case Study in Collaboration with UiPath Inc0
AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content0
A Retrospective Recount of Computer Architecture Research with a Data-Driven Study of Over Four Decades of ISCA Publications0
Automatic Knowledge Extraction with Human Interface0
Decontextualization: Making Sentences Stand-Alone0
Automated Parsing of Engineering Drawings for Structured Information Extraction Using a Fine-tuned Document Understanding Transformer0
DAViD: Domain Adaptive Visually-Rich Document Understanding with Synthetic Insights0
DavarOCR: A Toolbox for OCR and Multi-Modal Document Understanding0
Arctic-TILT. Business Document Understanding at Sub-Billion Scale0
Extract with Order for Coherent Multi-Document Summarization0
Auto-encodeurs pour la compr\'ehension de documents parl\'es (Auto-encoders for Spoken Document Understanding)0
A User-Centered Concept Mining System for Query and Document Understanding at Tencent0
CREPE: Coordinate-Aware End-to-End Document Parser0
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?0
ClueWeb22: 10 Billion Web Documents with Visual and Semantic Information0
Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration0
DrVideo: Document Retrieval Based Long Video Understanding0
Attention-Based Graph Neural Network with Global Context Awareness for Document Understanding0
Acronym Identification and Disambiguation Shared Tasks for Scientific Document Understanding0
ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding0
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling0
Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding0
A Multi-Modal Multilingual Benchmark for Document Image Classification0
DONUT-hole: DONUT Sparsification by Harnessing Knowledge and Optimizing Learning Efficiency0
A LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents0
DUBLIN -- Document Understanding By Language-Image Network0
Efficient End-to-End Visual Document Understanding with Rationale Distillation0
DOGE: Towards Versatile Visual Document Grounding and Referring0
Show:102550
← PrevPage 2 of 7Next →

No leaderboard results yet.