OCRmyPDF/README.md

<img src="docs/images/logo.svg" width="240" alt="OCRmyPDF">

[![Build Status][azure]](https://dev.azure.com/jim0585/ocrmypdf/_build/latest?definitionId=2&branchName=master) [![PyPI version][pypi]](https://pypi.org/project/ocrmypdf/) ![Homebrew version][homebrew] ![ReadTheDocs][docs] ![Python versions][pyversions]

[azure]: https://dev.azure.com/jim0585/ocrmypdf/_apis/build/status/jbarlow83.OCRmyPDF?branchName=master
[travis]: https://travis-ci.org/jbarlow83/OCRmyPDF.svg?branch=master "Travis build status"
[pypi]: https://img.shields.io/pypi/v/ocrmypdf.svg "PyPI version"
[homebrew]: https://img.shields.io/homebrew/v/ocrmypdf.svg "Homebrew version"
[docs]: https://readthedocs.org/projects/ocrmypdf/badge/?version=latest "RTD"
[pyversions]: https://img.shields.io/pypi/pyversions/ocrmypdf "Supported Python versions"

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted.

```bash
ocrmypdf                      # it's a scriptable command line program
   -l eng+fra                 # it supports multiple languages
   --rotate-pages             # it can fix pages that are misrotated
   --deskew                   # it can deskew crooked PDFs!
   --title "My PDF"           # it can change output metadata
   --jobs 4                   # it uses multiple cores by default
   --output-type pdfa         # it produces PDF/A by default
   input_scanned.pdf          # takes PDF input (or images)
   output_searchable.pdf      # produces validated PDF output
```

[See the release notes for details on the latest changes](https://ocrmypdf.readthedocs.io/en/latest/release_notes.html).

## Main features

- Generates a searchable [PDF/A](https://en.wikipedia.org/?title=PDF/A) file from a regular PDF
- Places OCR text accurately below the image to ease copy / paste
- Keeps the exact resolution of the original embedded images
- When possible, inserts OCR information as a "lossless" operation without disrupting any other content
- Optimizes PDF images, often producing files smaller than the input file
- If requested, deskews and/or cleans the image before performing OCR
- Validates input and output files
- Distributes work across all available CPU cores
- Uses [Tesseract OCR](https://github.com/tesseract-ocr/tesseract) engine to recognize more than [100 languages](https://github.com/tesseract-ocr/tessdata)
- Scales properly to handle files with thousands of pages
- Battle-tested on millions of PDFs

For details: please consult the [documentation](https://ocrmypdf.readthedocs.io/en/latest/).

## Motivation

I searched the web for a free command line tool to OCR PDF files: I found many, but none of them were really satisfying:

- Either they produced PDF files with misplaced text under the image (making copy/paste impossible)
- Or they did not handle accents and multilingual characters
- Or they changed the resolution of the embedded images
- Or they generated ridiculously large PDF files
- Or they crashed when trying to OCR
- Or they did not produce valid PDF files
- On top of that none of them produced PDF/A files (format dedicated for long time storage)

...so I decided to develop my own tool.

## Installation

Linux, Windows, macOS and FreeBSD are supported. Docker images are also available.

Users of Debian 9 or later or Ubuntu 16.10 or later may simply

```bash
apt-get install ocrmypdf
```

and users of Fedora 29 or later may simply

```bash
dnf install ocrmypdf
```

and Homebrew users (macOS, Linux, Windows Subsystem for Linux) may simply

```bash
brew install ocrmypdf
```

For everyone else, [see our documentation](https://ocrmypdf.readthedocs.io/en/latest/installation.html) for installation steps.

## Languages

OCRmyPDF uses Tesseract for OCR, and relies on its language packs. For Linux users, you can often find packages that provide language packs:

```bash
# Display a list of all Tesseract language packs
apt-cache search tesseract-ocr

# Debian/Ubuntu users
apt-get install tesseract-ocr-chi-sim  # Example: Install Chinese Simplified language pack

# Arch Linux users
pacman -S tesseract-data-eng tesseract-data-deu # Example: Install the English and German language packs
```

You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested.

## Documentation and support

Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via:

```bash
ocrmypdf --help
```

Our [documentation is served on Read the Docs](https://ocrmypdf.readthedocs.io/en/latest/index.html).

Please report issues on our [GitHub issues](https://github.com/jbarlow83/OCRmyPDF/issues) page, and follow the issue template for quick response.

## Requirements

In addition to the required Python version (3.6+), OCRmyPDF requires external program installations of Ghostscript, Tesseract OCR, QPDF, and Leptonica. OCRmyPDF is pure Python, but uses CFFI to portably generate library bindings. OCRmyPDF works on pretty much everything: Linux, macOS, Windows and FreeBSD.

## Press & Media

- [Going paperless with OCRmyPDF](https://medium.com/@ikirichenko/going-paperless-with-ocrmypdf-e2f36143f46a)
- [Converting a scanned document into a compressed searchable PDF with redactions](https://medium.com/@treyharris/converting-a-scanned-document-into-a-compressed-searchable-pdf-with-redactions-63f61c34fe4c)
- [c't 1-2014, page 59](https://heise.de/-2279695): Detailed presentation of OCRmyPDF v1.0 in the leading German IT magazine c't
- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670)
- [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html)
- [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/)

## Business enquiries

OCRmyPDF would not be the software that it is today without companies and users choosing to provide support for feature development and consulting enquiries. We are happy to discuss all enquiries, whether for extending the existing feature set, or integrating OCRmyPDF into a larger system.

## License

The OCRmyPDF software is licensed under the Mozilla Public License 2.0
(MPL-2.0). This license permits integration of OCRmyPDF with other code,
included commercial and closed source, but asks you to publish source-level
modifications you make to OCRmyPDF.

Some components of OCRmyPDF have other licenses, as noted in those files and the
``debian/copyright`` file. Most files in ``misc/`` use the MIT license, and the
documentation and test files are generally licensed under Creative Commons
ShareAlike 4.0 (CC-BY-SA 4.0).

## Disclaimer

The software is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
Add a logo 2019-07-07 01:07:48 -07:00			`<img src="docs/images/logo.svg" width="240" alt="OCRmyPDF">`
Change to README.md 2018-09-19 21:01:24 -07:00
v9.2.0 release notes and docs 2019-12-11 13:13:51 -08:00			`[![Build Status][azure]](https://dev.azure.com/jim0585/ocrmypdf/_build/latest?definitionId=2&branchName=master) [![PyPI version][pypi]](https://pypi.org/project/ocrmypdf/) ![Homebrew version][homebrew] ![ReadTheDocs][docs] ![Python versions][pyversions]`

			`[azure]: https://dev.azure.com/jim0585/ocrmypdf/_apis/build/status/jbarlow83.OCRmyPDF?branchName=master`
Fix broken badges in README 2018-10-12 21:16:08 -07:00			`[travis]: https://travis-ci.org/jbarlow83/OCRmyPDF.svg?branch=master "Travis build status"`
			`[pypi]: https://img.shields.io/pypi/v/ocrmypdf.svg "PyPI version"`
			`[homebrew]: https://img.shields.io/homebrew/v/ocrmypdf.svg "Homebrew version"`
Fix docs build 2018-11-12 13:26:04 -08:00			`[docs]: https://readthedocs.org/projects/ocrmypdf/badge/?version=latest "RTD"`
Python 3.8 updates 2019-10-20 03:20:54 -07:00			`[pyversions]: https://img.shields.io/pypi/pyversions/ocrmypdf "Supported Python versions"`

Change to README.md 2018-09-19 21:01:24 -07:00			`OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched or copy-pasted.`

			```bash
			`ocrmypdf # it's a scriptable command line program`
			`-l eng+fra # it supports multiple languages`
			`--rotate-pages # it can fix pages that are misrotated`
			`--deskew # it can deskew crooked PDFs!`
			`--title "My PDF" # it can change output metadata`
			`--jobs 4 # it uses multiple cores by default`
			`--output-type pdfa # it produces PDF/A by default`
			`input_scanned.pdf # takes PDF input (or images)`
			`output_searchable.pdf # produces validated PDF output`
			```

Bump pikepdf version, point to release notes 2019-01-05 16:48:13 -08:00			`[See the release notes for details on the latest changes](https://ocrmypdf.readthedocs.io/en/latest/release_notes.html).`

readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Main features`
Change to README.md 2018-09-19 21:01:24 -07:00
Readme: more media 2018-12-19 15:27:54 -08:00			`- Generates a searchable [PDF/A](https://en.wikipedia.org/?title=PDF/A) file from a regular PDF`
			`- Places OCR text accurately below the image to ease copy / paste`
			`- Keeps the exact resolution of the original embedded images`
			`- When possible, inserts OCR information as a "lossless" operation without disrupting any other content`
			`- Optimizes PDF images, often producing files smaller than the input file`
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`- If requested, deskews and/or cleans the image before performing OCR`
Readme: more media 2018-12-19 15:27:54 -08:00			`- Validates input and output files`
			`- Distributes work across all available CPU cores`
readme: tweaks 2019-03-16 14:09:19 -07:00			`- Uses [Tesseract OCR](https://github.com/tesseract-ocr/tesseract) engine to recognize more than [100 languages](https://github.com/tesseract-ocr/tessdata)`
			`- Scales properly to handle files with thousands of pages`
			`- Battle-tested on millions of PDFs`
Change to README.md 2018-09-19 21:01:24 -07:00
			`For details: please consult the [documentation](https://ocrmypdf.readthedocs.io/en/latest/).`

readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Motivation`
Change to README.md 2018-09-19 21:01:24 -07:00
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`I searched the web for a free command line tool to OCR PDF files: I found many, but none of them were really satisfying:`
Change to README.md 2018-09-19 21:01:24 -07:00
Readme: more media 2018-12-19 15:27:54 -08:00			`- Either they produced PDF files with misplaced text under the image (making copy/paste impossible)`
			`- Or they did not handle accents and multilingual characters`
			`- Or they changed the resolution of the embedded images`
			`- Or they generated ridiculously large PDF files`
			`- Or they crashed when trying to OCR`
			`- Or they did not produce valid PDF files`
			`- On top of that none of them produced PDF/A files (format dedicated for long time storage)`
Change to README.md 2018-09-19 21:01:24 -07:00
			`...so I decided to develop my own tool.`

readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Installation`
Change to README.md 2018-09-19 21:01:24 -07:00
v9.2.0 release notes and docs 2019-12-11 13:13:51 -08:00			`Linux, Windows, macOS and FreeBSD are supported. Docker images are also available.`
Change to README.md 2018-09-19 21:01:24 -07:00
			`Users of Debian 9 or later or Ubuntu 16.10 or later may simply`

			```bash
			`apt-get install ocrmypdf`
			```

Add Fedora install instructions. (#304) * Add Fedora install instructions. * Fix path to fedora_rawhide badget 2018-10-14 16:28:50 -04:00			`and users of Fedora 29 or later may simply`

			```bash
			`dnf install ocrmypdf`
			```

Generally update documentation about available platforms 2019-12-19 00:27:37 -08:00			`and Homebrew users (macOS, Linux, Windows Subsystem for Linux) may simply`
Change to README.md 2018-09-19 21:01:24 -07:00
			```bash
			`brew install ocrmypdf`
			```

			`For everyone else, [see our documentation](https://ocrmypdf.readthedocs.io/en/latest/installation.html) for installation steps.`

readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Languages`
Change to README.md 2018-09-19 21:01:24 -07:00
			`OCRmyPDF uses Tesseract for OCR, and relies on its language packs. For Linux users, you can often find packages that provide language packs:`

			```bash
			`# Display a list of all Tesseract language packs`
			`apt-cache search tesseract-ocr`

			`# Debian/Ubuntu users`
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`apt-get install tesseract-ocr-chi-sim # Example: Install Chinese Simplified language pack`

			`# Arch Linux users`
			`pacman -S tesseract-data-eng tesseract-data-deu # Example: Install the English and German language packs`
Change to README.md 2018-09-19 21:01:24 -07:00			```

			You can then pass the `-l LANG` argument to OCRmyPDF to give a hint as to what languages it should search for. Multiple languages can be requested.

readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Documentation and support`
Change to README.md 2018-09-19 21:01:24 -07:00
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`Once OCRmyPDF is installed, the built-in help which explains the command syntax and options can be accessed via:`
Change to README.md 2018-09-19 21:01:24 -07:00
			```bash
			`ocrmypdf --help`
			```

			`Our [documentation is served on Read the Docs](https://ocrmypdf.readthedocs.io/en/latest/index.html).`

Generally update documentation about available platforms 2019-12-19 00:27:37 -08:00			`Please report issues on our [GitHub issues](https://github.com/jbarlow83/OCRmyPDF/issues) page, and follow the issue template for quick response.`
Change to README.md 2018-09-19 21:01:24 -07:00
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Requirements`
Change to README.md 2018-09-19 21:01:24 -07:00
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`In addition to the required Python version (3.6+), OCRmyPDF requires external program installations of Ghostscript, Tesseract OCR, QPDF, and Leptonica. OCRmyPDF is pure Python, but uses CFFI to portably generate library bindings. OCRmyPDF works on pretty much everything: Linux, macOS, Windows and FreeBSD.`
Change to README.md 2018-09-19 21:01:24 -07:00
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Press & Media`
Change to README.md 2018-09-19 21:01:24 -07:00
Readme: more media 2018-12-19 15:27:54 -08:00			`- [Going paperless with OCRmyPDF](https://medium.com/@ikirichenko/going-paperless-with-ocrmypdf-e2f36143f46a)`
			`- [Converting a scanned document into a compressed searchable PDF with redactions](https://medium.com/@treyharris/converting-a-scanned-document-into-a-compressed-searchable-pdf-with-redactions-63f61c34fe4c)`
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`- [c't 1-2014, page 59](https://heise.de/-2279695): Detailed presentation of OCRmyPDF v1.0 in the leading German IT magazine c't`
			`- [heise Open Source, 09/2014: Texterkennung mit OCRmyPDF](https://heise.de/-2356670)`
Readme: Add another heise article 2020-02-18 02:41:22 -08:00			`- [heise Durchsuchbare PDF-Dokumente mit OCRmyPDF erstellen](https://www.heise.de/ratgeber/Durchsuchbare-PDF-Dokumente-mit-OCRmyPDF-erstellen-4607592.html)`
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`- [Excellent Utilities: OCRmyPDF](https://www.linuxlinks.com/excellent-utilities-ocrmypdf-add-ocr-text-layer-scanned-pdfs/)`
Change to README.md 2018-09-19 21:01:24 -07:00
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Business enquiries`
readme: tweaks 2019-03-16 14:09:19 -07:00
Fix typos, add instructions for training data (#477) 2020-01-31 01:24:41 +01:00			`OCRmyPDF would not be the software that it is today without companies and users choosing to provide support for feature development and consulting enquiries. We are happy to discuss all enquiries, whether for extending the existing feature set, or integrating OCRmyPDF into a larger system.`
readme: tweaks 2019-03-16 14:09:19 -07:00
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## License`
Change to README.md 2018-09-19 21:01:24 -07:00
Change license of all GPLv3 files to MPL-2.0 https://github.com/jbarlow83/OCRmyPDF/issues/600 2020-08-05 00:44:42 -07:00			`The OCRmyPDF software is licensed under the Mozilla Public License 2.0`
			`(MPL-2.0). This license permits integration of OCRmyPDF with other code,`
			`included commercial and closed source, but asks you to publish source-level`
			`modifications you make to OCRmyPDF.`

			`Some components of OCRmyPDF have other licenses, as noted in those files and the`
			``debian/copyright`` file. Most files in ``misc/`` use the MIT license, and the
			`documentation and test files are generally licensed under Creative Commons`
			`ShareAlike 4.0 (CC-BY-SA 4.0).`
Change to README.md 2018-09-19 21:01:24 -07:00
readme: markdown cleanup 2020-06-29 01:45:27 -07:00			`## Disclaimer`
Change to README.md 2018-09-19 21:01:24 -07:00
			`The software is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.`