SOTAVerified

Optical Character Recognition (OCR)

Optical Character Recognition or Optical Character Reader (OCR) is the electronic or mechanical conversion of images of typed, handwritten or printed text into machine-encoded text, whether from a scanned document, a photo of a document, a scene-photo (for example the text on signs and billboards in a landscape photo, license plates in cars...) or from subtitle text superimposed on an image (for example: from a television broadcast)

Papers

Showing 401450 of 1209 papers

TitleStatusHype
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsCode2
Constructing Image-Text Pair Dataset from Books0
Comprehensive Overview of Named Entity Recognition: Models, Domain-Specific Applications and Challenges0
Order-preserving Consistency Regularization for Domain Adaptation and GeneralizationCode0
STEP -- Towards Structured Scene-Text SpottingCode0
Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach0
Separate and Locate: Rethink the Text in Text-based Visual Question AnsweringCode0
DTrOCR: Decoder-only Transformer for Optical Character RecognitionCode2
Enhancing OCR Performance through Post-OCR Models: Adopting Glyph Embedding for Improved Correction0
Vision Grid Transformer for Document Layout Analysis0
Optimal Projections for Discriminative Dictionary Learning using the JL-lemmaCode0
Bengali Document Layout Analysis with Detectron20
DISGO: Automatic End-to-End Evaluation for Scene Text OCR0
Nougat: Neural Optical Understanding for Academic DocumentsCode5
American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers0
CNN based Cuneiform Sign Detection Learned from Annotated 3D Renderings and Mapped Photographs with Illumination Augmentation0
bbOCR: An Open-source Multi-domain OCR Pipeline for Bengali DocumentsCode1
BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual QuestionsCode2
OCR Language Models with Custom Vocabularies0
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo EmbeddingsCode0
OmniDataComposer: A Unified Data Structure for Multimodal Data Fusion and Infinite Data GenerationCode1
Training BERT Models to Carry Over a Coding System Developed on One Corpus to Another0
Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character RecognitionCode1
Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level Annotations0
Making the V in Text-VQA Matter0
Optimizing the Neural Network Training for OCR Error Correction of Historical Hebrew Texts0
Toward a Period-Specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts0
Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math TextbooksCode0
Multi-Granularity Prediction with Learnable Fusion for Scene Text Recognition0
MataDoc: Margin and Text Aware Document Dewarping for Arbitrary Boundary0
A comparative analysis of SRGAN models0
Modular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific Documents (Extended Version)Code1
Handwritten and Printed Text Segmentation: A Signature Case Study0
Handwritten Text Recognition Using Convolutional Neural Network0
A Novel Pipeline for Improving Optical Character Recognition through Post-processing Using Natural Language Processing0
Artificial Eye for the Blind0
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding0
Estimating Post-OCR Denoising Complexity on Numerical Texts0
Fraunhofer SIT at CheckThat! 2023: Mixing Single-Modal Classifiers to Estimate the Check-Worthiness of Multi-Modal Tweets0
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image UnderstandingCode2
UTRNet: High-Resolution Urdu Text Recognition In Printed DocumentsCode1
Resume Information Extraction via Post-OCR Text Processing0
A Survey on Multimodal Large Language Models0
Document Image Cleaning using Budget-Aware Black-Box ApproximationCode0
GenPlot: Increasing the Scale and Diversity of Chart Derendering DataCode1
Weakly supervised information extraction from inscrutable handwritten document images0
When Vision Fails: Text Attacks Against ViT and OCRCode0
SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure CaptioningCode0
Transformer-Based UNet with Multi-Headed Cross-Attention Skip Connections to Eliminate Artifacts in Scanned Documents0
TransDocAnalyser: A Framework for Offline Semi-structured Handwritten Document Analysis in the Legal DomainCode1
Show:102550
← PrevPage 9 of 25Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1DTrOCR 105MAccuracy (%)89.6Unverified
2DTrOCRAccuracy (%)89.6Unverified
3MaskOCR-LAccuracy (%)82.6Unverified
4TransOCRAccuracy (%)72.8Unverified
5SRNAccuracy (%)65Unverified
6MORANAccuracy (%)64.3Unverified
7SEEDAccuracy (%)61.2Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4oAverage Accuracy76.22Unverified
2Gemini-1.5 ProAverage Accuracy76.13Unverified
3Claude-3 SonnetAverage Accuracy67.71Unverified
4RapidOCRAverage Accuracy56.98Unverified
5EasyOCRAverage Accuracy49.3Unverified
#ModelMetricClaimedVerifiedStatus
1STREETSequence error27.54Unverified
2SEESequence error22Unverified
3AttentionOCR_Inception-resnet-v2_LocationSequence error15.8Unverified
#ModelMetricClaimedVerifiedStatus
1I2L-NOPOOLBLEU89.09Unverified
2I2L-STRIPSBLEU89Unverified
#ModelMetricClaimedVerifiedStatus
1TesseractCharacter Error Rate (CER)0.08Unverified
2EasyOCRCharacter Error Rate (CER)0.07Unverified
#ModelMetricClaimedVerifiedStatus
1I2L-STRIPSBLEU88.86Unverified