Release date:
The pdfOCR add-on for iText Core enables OCR processing for scanned PDFs and images, making text searchable and extractable. As well as Tesseract, it also supports ONNX-based recognition engines with optional GPU acceleration.
There are no feature changes for this release. The only changes are to maintain compatibility with the iText Core 9.8.0 dependencies.
Downloads
|
|
||||
|---|---|---|---|---|
|
iText pdfOCR – 5.0.2 (Java) |
link (API) link (Tesseract) link (ONNX-abstract) link (ONNX-cpu) |
N/A |
link (API) link (Tesseract) link (ONNX-abstract) link (ONNX-cpu) |
|
|
iText pdfOCR – 5.0.2 (.NET) |
N/A |
link (API) link (Tesseract) link (ONNX-abstract) link (ONNX-cpu) |
link (API) link (Tesseract) link (ONNX-abstract) link (ONNX-cpu) |
Installation Instructions
Examples (latest ones)
FAQ (latest ones)
- pdfOCR: Who provided TESS_DATA_DIRECTORY?
- pdfOCR: Is handwriting recognition supported
- pdfOCR: If your scanned document has a mixture of sections with paragraphs and tables, what is a recommended strategy here?
- How do I create a separate OCR layer?
- Could not find a glyph corresponding to Unicode character