DvD
(paper:2025) - Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
An awesome list with 278 entries.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
(paper:2025) - Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
(paper:2025) - Geometry Restoration and Dewarping of Camera-Captured Document Images
(paper:2022) - Adaptive Radial Projection on Fourier Magnitude Spectrum for Document Image Skew Estimation
(paper:2019)
(paper:2019) - A Multi-Object Rectified Attention Network for Scene Text Recognition
(2020) An application of high resolution GANs to dewarp images of perturbed documents
(2016) - Perspective recovery of text using transformed ellipses
a post-processing tool for scanned sheets of paper, especially for book pages that have been scanned from previously created photocopies.
Library used to deskew a scanned document
Contains code to deskew images using MLPs, LSTMs and LLS tranformations
De-skewing images with slanted content by finding the deviation using Canny Edge Detection.
(2016) - Page dewarping and thresholding using a "cubic sheet" model
Rotate text images if they are not straight for better text detection and recognition.
Deskew is a command line tool for deskewing scanned text documents. It uses Hough transform to detect "text lines" in the image. As an output, you get an image rotated so that the lines are horizontal.
No code :(
Document Image Dewarping Using Text Lines and Line Segments via Classical Methods Only (Without Machine Learning)
Deep Learning Chinese Word Segment
Detect handwritten words (classic image processing based method).
LAREX is a semi-automatic open-source tool for layout analysis on early printed books.
Deep learning based page layout analysis
Page to PAGE Layout Analysis Tool
This is a deep learning model for page layout analysis / segmentation.
Layout Analysis Evaluator for the ICDAR 2017 competition on Layout Analysis for Challenging Medieval Manuscripts
a deep learning model for page layout analysis / segmentation.
Deep Learning Chinese Word Segment
Handwritten Text Recognition (HTR) system implemented with TensorFlow.
OCR software for recognition of handwritten text
apply HTR services from Amazon, Google, and/or Microsoft to scanned documents
paper:2024 UniTable: Towards a Unified Table Foundation Model
Unofficial implementation of ICDAR 2019 paper : TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images.
Table Extraction Tool
Table recognition inside douments using neural networks.
Large-scale table detection and recognition dataset with pre-trained models
Extract tables from scanned image PDFs using Optical Character Recognition.
The most accurate natural language detection library for Java and other JVM languages, suitable for long and short text alike
Lightning Fast Language Prediction rocket
Text language identification using Wikipedia data. No license specified.
paper:2018) - Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation
(paper:2018) - Arbitrary-Oriented Scene Text Detection via Rotation Proposals
(paper:2021) - TensorFlow reimplementation of "MASTER: Multi-Aspect Non-local Network for Scene Text Recognition" (Pattern Recognition 2021).
(paper:2020) - Mask TextSpotter v3 is an end-to-end trainable scene text spotter that adopts a Segmentation Proposal Network (SPN) instead of an RPN.
(paper:2020) A PyTorch implementation of "TextFuseNet: Scene Text Detection with Richer Fused Features".
(paper:2020) - Official Tensorflow Implementation of Self-Attention Text Recognition Network (SATRN) (CVPR Workshop WTDDLE 2020).
(paper:2020) - Unofficial implementation of CVPR 2020 paper "SCATTER: Selective Context Attentional Scene Text Recognizer"
([paper:2020[https://arxiv.org/pdf/2005.10977.pdf]) - This is the implementation of the paper "SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition"
A scene text recognition toolbox based on PyTorch
(paper:2020) Efficient Backbone Search for Scene Text Recognition
(paper:2019) Pytorch implementation for "Decoupled attention network for text recognition".
(paper:2020) Implementation of Bidirectional Scene Text Recognition with a Single Decoder
(paper:2019
(paper:2019)
(paper:2019)
(paper:2019
(paper:2019)
(paper:2019)
Repository for Scene Text Detection with Supervised Pyramid Context Network with tensorflow.
End-to-end pipeline for real-time scene text detection and recognition.
Robust Scene Text Recognition with Automatic Rectification.
The code of "Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes".
An Implementation of the alogrithm in paper IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection.
An End-to-End TextSpotter with Explicit Alignment and Attention
RRD: Rotation-Sensitive Regression for Oriented Scene Text Detection.
Implement 'Single Shot Text Detector with Regional Attention, ICCV 2017 Spotlight'.
caffe re-implementation of R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection.
This project modify tensorflow object detection api code to predict oriented bounding boxes. It can be used for scene text detection.
This is a c++ project deploying a deep scene text reading pipeline with tensorflow. It reads text from natural scene images. It uses frozen tensorflow graphs. The detector detect scene text locations. The recognizer reads word from each detected bounding box.
A novel region proposal network for more general object detection ( including scene text detection ).
SEE: Towards Semi-Supervised End-to-End Scene Text Recognition
Code for the paper STN-OCR: A single Neural Network for Text Detection and Text Recognition
Scene text detection and recognition based on Extremal Region(ER)
Rotational region detection based on Faster-RCNN.
Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation.
A PyTorch implementation of ECCV2018 Paper: TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes
Implementation for CVPR 2018 text recognition Paper by Tensorflow: "AON: Towards Arbitrarily-Oriented Text Recognition"
Implementation of our paper 'PixelLink: Detecting Scene Text via Instance Segmentation' in AAAI2018
An Implementation of the seglink alogrithm in paper Detecting Oriented Text in Natural Images by Linking Segments (=> pixe_link)
Single Shot Text Detector with Regional Attention
(paper:2019) - A Multi-Object Rectified Attention Network for Scene Text Recognition
This repository provides train&test code, dataset, det.&rec. annotation, evaluation script, annotation tool, and ranking table.
TextField: Learning A Deep Direction Field for Irregular Scene Text Detection (TIP 2019)
TextMountain: Accurate Scene Text Detection via Instance Segmentation
Recognizing cropped text in natural images.
A fuzzy receipt parser written in Python.
Pytorch implementation of CRAFT text detector.
Text recognition combo - CRAFT + CRNN.
PyTorch implementation of CRAFT
A packaged and flexible version of the CRAFT text detector and Keras CRNN recognition model.
TextBoxes++: A Single-Shot Oriented Scene Text Detector
Textboxes_plusplus implementation with Tensorflow (python)
This is a tensorflow re-implementation of PSENet: Shape Robust Text Detection with Progressive Scale Expansion Network
Shape Robust Text Detection with Progressive Scale Expansion Network.
(official) - (tf1/py2) A tensorflow implementation of EAST text detector
(tf1/py2) AdvancedEAST is an algorithm used for Scene image text detect, which is primarily based on EAST, and the significant improvement was also made, which make long text predictions more accurate.
Implementation of EAST scene text detector in Keras
This is a pytorch re-implementation of EAST: An Efficient and Accurate Scene Text Detector.
Forked from argman/EAST for the ICPR MTWI 2018 CHALLENGE
Text Detection from images using OpenCV
TextBoxes re-implement using tensorflow
TextBoxes: A Fast Text Detector with a Single Deep Neural Network
Textboxes implementation with Tensorflow (python)
Textboxes : Image Text Detection Model : python package (tensorflow)
Detecting Text in Natural Image with Connectionist Text Proposal Network
The first open-source library that detects the font of a text in a image.
Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction
(paper:2026) - Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
OCR model for math that outputs LaTeX and markdown.
(paper:2021)
(paper:2021) - TensorFlow reimplementation of "MASTER: Multi-Aspect Non-local Network for Scene Text Recognition" (Pattern Recognition 2021).
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Ocular is a state-of-the-art historical OCR system.
python ocr using tesseract/ with EAST opencv detector
Tensorflow-based CNN+LSTM trained with CTC-loss for OCR.
Scan the MRZ code of a passport and extract the firstname, lastname, passport number, nationality, date of birth, expiration date and personal numer.
OCR using tensorflow with attention.
OCR detection implement with tensorflow v1.4.
This is the implementation of the paper "Gated Recurrent Convolution Neural Network for OCR"
A tool for extracting text from scanned documents (via OCR), with user-defined post-processing.
MXNet OCR implementation. Including text recognition and detection.
The first Xi'an Jiaotong University Artificial Intelligence Practice Contest (2018AI Practice Contest - Picture Text Recognition) first; only use the densenet to identify the Chinese characters
CNN+LSTM+CTC based OCR implemented using tensorflow.
A small C++ implementation of LSTM networks, focused on OCR.
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
Tesseract Open Source OCR Engine (main repository)
Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.
Repository collecting all the submodules for the new PyTorch-based OCR System.
Python-based tools for document analysis and OCR
make a better chinese character recognition OCR than tesseract.
Library for efficient text classification and representation learning
A Tensorflow model for text recognition (CNN + seq2seq with visual attention) available as a Python package and compatible with Google Cloud ML Engine.
Visual Attention based OCR
Table recognition inside douments using neural networks.
A machine learning software for extracting information from scholarly documents
LA-PDFText is a system for extracting accurate text from PDF-based research articles
LAREX is a semi-automatic open-source tool for layout analysis on early printed books.
Go package for OCR (Optical Character Recognition), by using Tesseract C++ library.
Automatically extract printed text, handwriting, and data from any document.
Simple app for visual editing of Page XML files
Validate and transform various OCR file formats
Tools for manipulating and evaluating the hOCR format for representing multi-lingual OCR results by embedding them into HTML.
Total Text Dataset. It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind.
This is a synthetically generated dataset, in which word instances are placed in natural scene images, while taking into account the scene layout.
DIAR software for synthetic document image and groundtruth generation, with various degradation models for data augmentation.
Scene Text Image Transformer
A synthetic data generator for text recognition
Pre-Recognize Library - library with algorithms for improving OCR quality.
Toolbox for OCR post-correction
The CIS OCR PostCorrectionTool
Text recognition (optical character recognition) with deep learning methods.
dinglehopper is an OCR evaluation tool and reads ALTO, PAGE and text files.
a small Python library implementing document image degradation for data augmentation for handwriting recognition and OCR applications.
Scan Tailor is an interactive post-processing tool for scanned pages.
help researchers fix these errors and extract the highest quality text from their pdfs as possible.
Simple app for visual editing of Page XML files
Open Semantic Search Engine and Open Source Text Mining & Text Analytics platform (Integrates ETL for document processing, OCR for images & PDF, named entity recognition for persons, organizations & locations, metadata management by thesaurus & ontologies, search user interface & search apps for…
A simple OCR API server, seriously easy to be deployed by Docker, on Heroku as well
~1000 book pages + OpenCV + python = page regions identified as paragraphs, lines, images, captions, etc.
An expandable and scalable OCR pipeline
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…