Provides OCR (Optical Character Recognition) services through web applications
10K+
Note
OCR4all is no longer under active development and has entered maintenance mode. The existing codebase remains available for continued use and reference, but no new features or major enhancements are planned. All future development efforts will be focused on [LAREX](https://github.com/OCR4all/larex)
As suggested by the name one of the main goals of OCR4all is to allow basically any given user to independently perform OCR on a wide variety of historical printings and obtain high quality results with reasonable time expenditure. Therefore, OCR4all is explicitly geared towards users with no technical background. If you are one of those users (or if you just want to use the tool and are not interested in the code), please go to the documentation website or the getting started project where you will find test data.
Please note that OCR4all current main focus is a semi-automatic workflow allowing users to perform OCR even on the earliest printed books, which is a very challenging task that often requires a significant amount of manual interaction, especially when almost perfect quality is desired. Nevertheless, we are working towards increasing robustness and the degree of automation of the tool. An important cornerstone for this is the recently agreed cooperation with the OCR-D project which focuses on the mass full-text recognition of historical materials.
This repository contains the code for the main interface and server of the OCR4all project, while the repositories OCR4all/docker_image and OCR4all/docker_base_image are about the creation of a preconfigurated docker image.
For installing the complete project with a docker image, please follow the instructions here.
OCR4all is under active development and consequently, frequent releases containing bug fixes and further functionality can be expected. In order to always be up to date, we highly recommend subscribing to our mailing list where we will always announce notable enhancements.
If you are using OCR4all please cite:
Reul, C., Christ, D., Hartelt, A., Balbach, N., Wehner, M., Springmann, U., Wick, C., Grundig, Büttner, A., C., Puppe, F.: OCR4all — An open-source tool providing a (semi-) automatic OCR workflow for historical printings Applied Sciences 9(22) (2019)
@article{reul2019ocr4all,
title={OCR4all—An open-source tool providing a (semi-) automatic OCR workflow for historical printings},
author={Reul, Christian and Christ, Dennis and Hartelt, Alexander and Balbach, Nico and Wehner, Maximilian and Springmann, Uwe and Wick, Christoph and Grundig, Christine and B{\"u}ttner, Andreas and Puppe, Frank},
journal={Applied Sciences},
volume={9},
number={22},
pages={4853},
year={2019},
publisher={Multidisciplinary Digital Publishing Institute}
}
Content type
Image
Digest
sha256:9074d4a85…
Size
8 GB
Last updated
over 2 years ago
docker pull uniwuezpd/ocr4all