PyPDFOCR on Docker
get rid of your paperwork...
PyPDFOCR converts a scanned PDF into an OCR'ed PDF using Tesseract-OCR and Ghostscript. When watching a folder it sorts the documents then automatically into folders defined with keywords in your config file.
This Docker image is based on the official Ubuntu base image.
It incorporates a patch for issue #41 of pypdfocr 0.9.0 likely to be fixed in 0.9.1
docker run --rm mmatiaschek/pypdfocr [-h] [-d] [-v] [-m] [-l LANG] [--preprocess]
[--skip-preprocess] [-w WATCH_DIR] [-f] [-c CONFIGFILE] [-e]
[-n]
[pdf_filename]
docker run -v ~/:/media --rm pypdfocr /media/filename.pdf
--> reads filename.pdf from your Home directory, filename_ocr.pdf will be generated
docker run -v ~/Documents/Paper:/media --rm mmatiaschek/pypdfocr -w /media -f -c /media/config.yaml
For sample config see config.yaml or pypdfocr authors repository here.
docker run --rm mmatiaschek/pypdfocr [-h] [-d] [-v] [-m] [-l LANG] [--preprocess]
[--skip-preprocess] [-w WATCH_DIR] [-f] [-c CONFIGFILE] [-e]
[-n]
[pdf_filename]
Interactive Shell
docker run --entrypoint=/bin/bash -t -i mmatiaschek/pypdfocr
This way my personal documents don't have to leave my hardware or network aka personal cloud.
Content type
Image
Digest
Size
275.9 MB
Last updated
almost 10 years ago
docker pull mmatiaschek/pypdfocr