SOTAVerified

Document AI

Papers

Showing 1–40 of 40 papers

TitleStatusHype
DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksCode4
Unifying Vision, Text, and Layout for Universal Document ProcessingCode3
LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingCode2
Document AI: A Comparative Study of Transformer-Based, Graph-Based Models, and Convolutional Neural Networks For Document Layout AnalysisCode1
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information ExtractionCode1
Document Intelligence Metrics for Visually Rich Document EvaluationCode1
Document Understanding Dataset and Evaluation (DUDE)Code1
Modular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific Documents (Extended Version)Code1
Context-Aware Chart Element DetectionCode1
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office AutomationCode1
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and ReasoningCode1
DiT: Self-supervised Pre-training for Document Image TransformerCode1
DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine ReadingCode1
LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingCode0
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document ClassificationCode0
DoSA : A System to Accelerate Annotations on Business Documents with Human-in-the-LoopCode0
Design of a Quality Management System based on the EU Artificial Intelligence ActCode0
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingCode0
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document ParsingCode0
DiMSum: Distributed and Multilingual Summarization of Financial NarrativesCode0
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form ParserCode0
H2OVL-Mississippi Vision Language Models Technical Report—0
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations—0
Can AI Models Appreciate Document Aesthetics? An Exploration of Legibility and Layout Quality in Relation to Prediction Confidence—0
Development of a Legal Document AI-Chatbot—0
Document AI: Benchmarks, Models and Applications—0
DocXChain: A Powerful Open-Source Toolchain for Document Parsing and Beyond—0
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment—0
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts—0
FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction—0
GeoLayoutLM: Geometric Pre-training for Visual Information Extraction—0
A Multi-Modal Multilingual Benchmark for Document Image Classification—0
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images—0
LongFin: A Multimodal Document Understanding Model for Long Financial Domain Documents—0
Model Reporting for Certifiable AI: A Proposal from Merging EU Regulation into AI Development—0
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding—0
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts—0
PrIeD-KIE: Towards Privacy Preserved Document Key Information Extraction—0
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents—0
Vision Grid Transformer for Document Layout Analysis—0
Show:102550

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1LayoutLMv3Average F199.21—Unverified