OCRmyPDF

mirror of https://github.com/ocrmypdf/OCRmyPDF.git synced 2025-10-24 22:39:05 +00:00

Author	SHA1	Message	Date
Jim Barlow	ec8a35a7a6	Basic test cases	2015-07-22 02:59:25 -07:00
Jim Barlow	f6577c22c3	Complete wrapping of logger/logger_mutex	2015-07-22 02:57:13 -07:00
Jim Barlow	43d6c03093	Implement oversampling in ocrpage.py	2015-03-27 18:32:55 -07:00
Jim Barlow	1870f116bb	More consistent spacing	2015-03-24 23:05:42 -07:00
Jim Barlow	8b87def013	Don't presume two jobs	2015-03-24 23:04:49 -07:00
Jim Barlow	de599d97b5	Tidy up readme	2015-03-24 23:04:33 -07:00
Jim Barlow	5d7e6b45c4	Cleanup logger	2015-03-24 22:46:33 -07:00
Jim Barlow	c6091bcfe1	Change python2 -> python3 for readlink()	2015-03-24 22:36:13 -07:00
Jim Barlow	466a8a1318	It's now py3 that uses lxml, reportlab	2015-03-19 17:12:32 -07:00
Jim Barlow	a99ba3b696	Add rudimentary support for combining OCR layer with existing content It appears to be very fragile due to weaknesses in PyPDF. Better option is probably to use pdftk's watermark feature.	2015-03-10 14:28:38 -07:00
Jim Barlow	9229f7c6cc	Add option to render text as invisible OCR text Prior to this change, hocrtransform would render printable text (black on white) and then a fully opaque image on top of the text. According to the PDF spec, text that is the output of OCR should be marked invisible, so that PDF viewers /know/ it's OCR output in a document that might mix OCR and text overlays. Another benefit is that PDF viewers would know to skip rendering text if they are not smart enough to figure out the image will completely overwrite it. However, for debug, visible text is nice, so retain it as an option.	2015-02-22 12:43:27 -08:00
Jim Barlow	bf114bb188	Clean up pixel transform logic with namedtuple	2015-02-21 14:14:34 -08:00
Jim Barlow	b8eed2f861	More PEP8/lint	2015-02-21 13:00:46 -08:00
Jim Barlow	ccb1e347be	Call HocrTransform directly instead of through a subprocess	2015-02-20 17:20:48 -08:00
Jim Barlow	8698974f11	Rename hocrTransform -> hocrtransform	2015-02-20 16:47:36 -08:00
Jim Barlow	f2c79c4341	Convert hocrtransform to py3	2015-02-20 16:38:24 -08:00
Jim Barlow	4966d1346b	Module marker for src folder	2015-02-20 15:43:05 -08:00
Jim Barlow	4a9337f757	PEP8	2015-02-20 15:42:06 -08:00
Jim Barlow	db311fb6a2	Add support for -b (skip big pages)	2015-02-20 15:26:33 -08:00
Jim Barlow	02c1dcec8e	Remove filenames from .hocr files As documented, Tesseract does not escape the filename when inserting it into .hocr, potentially creating an invalid XML file as a result. Since there is no use for the title, regex it and nuke it.	2015-02-13 13:41:14 -08:00
Jim Barlow	52dc74d3ce	Support Tesseract 3.03 quirk: .html vs .hocr extension	2015-02-11 10:24:10 -08:00
Jim Barlow	cc2af2bc15	Convert the final image to a JPEG if the original image was a JPEG Of course, this introduces recompression artifacts, and is unnecessary if no options are given that modify the final image (no -d, -c, -i). But rather than worry about that, it would be better to ultimately find a way to combine the original PDF page with the output PDF text in the case where we want no changes to the original. This is good enough for now. The better option can apparently be achieved using pdftk background, or probably better, PyPDF2's merge. If Tesseract PDF generation is used then we need a way to remove the image. Tesseract PDF generation at 3.03 does layout better (I think) and also properly encodes the hidden layer, which is less likely to give display issues (I think).	2015-02-11 10:23:45 -08:00
Jim Barlow	638c6db05d	Use the appropriate PNG rendered given the types of image present	2015-02-11 03:32:00 -08:00
Jim Barlow	f7db8d9aff	Use Ghostscript -> PNG instead of pdftoppm for rendering Ghostscript has the clunkiest imaginable syntax, obtuse documentation, quirky behavior, and poor diagnostics... but it actually works unlike pdftoppm/poppler which gets things wrong. In this case I observed poppler incorrectly decompresses certain CCITT encoded monochrome PDFs. So set up Ghostscript to do the job instead. For the moment this performs monochrome -> RGB conversion via reportlab.	2015-02-11 03:13:07 -08:00
Jim Barlow	564fb7a87e	Support Ghostscript 9.14's new color conversion engine (not portable) The flag -dUseCIEColor is now deprecated, as it invokes the old engine which introduces color errors. The new engine requires a PDF/A file header with hardcoded location of a ICC profile to use, now included in the project. Portable iterations should generate a PDFA_def.ps based on the target system; for now OS X with homebrew is presumed. I have selected sRGB since scanners tend to capture RGB and printing is not a major consideration for PDF/A. Also note all file paths given to gs must be absolute. May its creators be forever haunted for their failure to document this unexpected quirk.	2015-02-09 15:33:49 -08:00
Jim Barlow	4d88e64774	Standardize tmpfile prefix	2015-02-09 15:02:49 -08:00
Jim Barlow	26f1163b46	Handle case where a page contains no images - don't OCR It doesn't make much sense to do anything with an all vector page except extract the page unmodified.	2015-02-08 20:05:54 -08:00
Jim Barlow	40058e99e0	Implement debug text only page option	2015-02-08 19:51:41 -08:00
Jim Barlow	bece4c3e02	Describe what decision was made based on -f and -s and presence of text	2015-02-08 19:51:18 -08:00
Jim Barlow	f0f6b57c87	When deciding on OCR, check for presence of text rather than a font It appears to be possible to have a PDF with an embedded font that is either unused or used only for whitespace. So check for some amount of actual text instead.	2015-02-08 17:38:27 -08:00
Jim Barlow	dc2a4ab044	Logic error	2015-02-08 17:33:35 -08:00
Jim Barlow	b16d6f5b81	Implement skipping OCR when -s is specified Appears to be necessary to disable each state of the pipeline that is inactive, not just initial and terminal stages of an inactive segment. If nothing else this makes what is going on more explicit.	2015-02-08 17:26:16 -08:00
Jim Barlow	69ce6ff7b5	Not a named param	2014-11-22 15:35:05 -08:00
Jim Barlow	32ba50b8dc	Add Tesseract timeout to keep things reasonable	2014-11-14 02:06:23 -08:00
Jim Barlow	36aca45f35	The -dci options now work (and valid combinations thereof)	2014-11-14 00:23:22 -08:00
Jim Barlow	925290342d	Leptonica deskew can handle .pnm input, unlike imagemagick	2014-11-13 23:20:25 -08:00
Jim Barlow	4dc0370c57	Add leptonica deskew	2014-11-13 16:53:26 -08:00
Jim Barlow	b92f8e43f2	Run as a module instead	2014-11-13 16:52:53 -08:00
Jim Barlow	22b0733a1d	Merge branch 'feature/findskew' into develop	2014-11-13 16:00:27 -08:00
Jim Barlow	6021684ab6	Attempt to fix multiprocessing pickling error	2014-11-13 15:58:57 -08:00
Jim Barlow	f4b1d0cdfe	Fix symlink error that occurs in multipage processing	2014-11-13 15:58:36 -08:00
Jim Barlow	d0d8048621	Comments	2014-10-17 17:28:31 -07:00
Jim Barlow	cfd119325d	Use abspath instead of relpath for temporary directory symlink	2014-10-11 17:48:56 -07:00
Jim Barlow	ad30833ffc	Support missing tess_cfg_files parameter when omitted by OCRmyPDF.sh	2014-10-11 17:48:33 -07:00
Jim Barlow	e5c79a6666	Use TIFFs as intermediates pdftoppm in recent versions (0.26.4,5) seems to be incapable of producing valid TIFFs, so have it dump a .pnm file and let ImageMagick figure out how to convert it to TIFF. This is not ideal, but at least it works.	2014-10-10 01:54:16 -07:00
Jim Barlow	63dc753c1b	Standardize intermediate filenames better convert .pnm -deskew <...> .pnm seems to have a bug that produces an invalid .pnm file which later causes tesseract (specifically, leptonica) to choke (using 3.02/1.71 as versions, respectively). Will change pipeline to use tiffs internally since they are less stupid.	2014-10-10 01:30:43 -07:00
Jim Barlow	017bc1f252	Basic error handling	2014-10-10 01:07:46 -07:00
Jim Barlow	bcd67c009d	Sort of working, but fragile; uses tmp folder properly now	2014-10-10 00:35:49 -07:00
fritz-hh	635358884e	start rewrite ocrmypdf in python	2014-10-09 22:53:08 +02:00
Jim Barlow	2f6cfafdfc	Now produces a finished OCR-PDF page	2014-10-08 03:54:06 -07:00

... 51 52 53 54 55 ...

2895 Commits