Smart, Modern, and Extensible PDF-to-Text Converter with Advanced OCR and GUI
PDFTextGenie is a next-generation, user-friendly Python application for extracting clean, accurate text from scanned or image-based PDFs. It leverages state-of-the-art OCR (Optical Character Recognition) engines and advanced post-processing to deliver high-fidelity text output, even from noisy or complex documents. Designed for researchers, students, archivists, and anyone who needs reliable text extraction, PDFTextGenie offers batch processing, smart formatting, and a modern, cross-platform GUI.
- Batch PDF Input: Select and queue multiple PDFs for conversion.
- Advanced OCR: Uses EasyOCR (with modular backend for future engines like TrOCR, Tesseract, LLaMA OCR).
- Image Preprocessing: Binarization, denoising, and contrast enhancement for better OCR accuracy.
- Smart Text Cleanup: Removes extra newlines, fixes hyphenated line breaks, merges paragraphs.
- Modern GUI: Built with PyQt5, featuring drag & drop, live status, output preview, and persistent settings.
- Error Handling: User-friendly feedback for failed conversions, with detailed error dialogs.
- Persistent Settings: Remembers your last-used options and output folder.
- Extensible Architecture: Modular codebase for easy addition of new OCR engines or output formats.
- Offline & Cross-Platform: No API keys or internet required. Works on Windows, Linux, and macOS.
pdftextgenie/
├── gui/
│ └── main_window.py # PyQt5 GUI
├── ocr/
│ ├── llama_ocr.py # OCR backend (EasyOCR, extensible)
│ └── postprocess.py # Text cleanup utilities
├── utils/
│ └── pdf_utils.py # PDF to image conversion & preprocessing
├── output/
│ └── (Generated .txt files)
├── tests/
│ └── test_pdftextgenie.py # Automated tests
├── app.py # Main launcher
├── requirements.txt # Dependencies
└── README.md # This file
git clone https://github.com/-------
cd pdftextgeniepip install -r requirements.txtNote: You may need additional system packages for
pdf2image(e.g., poppler). See pdf2image docs for details.
python app.py- Use the GUI to select PDFs, set options, and start conversion.
- Extracted text files will appear in your chosen output folder.
- Add PDFs: Click 'Add PDFs' or drag & drop files into the list.
- Remove/Clear: Remove selected or clear all PDFs from the list.
- Choose Output Folder: Set where extracted text files will be saved.
- Options:
- Fix hyphens: Reconstruct words split across lines.
- Merge paragraphs: Combine lines into paragraphs.
- DPI: Set image resolution for PDF-to-image conversion.
- Image Preprocessing:
- Binarize: Convert images to black & white for better OCR.
- Denoise: Apply median filter to reduce noise.
- Enhance Contrast: Boost contrast (set factor).
- Start Conversion: Begin batch processing. Progress and errors are shown live.
- Preview: See a snippet of the extracted text and any errors.
- Help/About: Access app info from the menu bar.
- OCR Backend:
- Currently uses EasyOCR. The code is modular—add new engines (Tesseract, TrOCR, etc.) in
ocr/llama_ocr.py.
- Currently uses EasyOCR. The code is modular—add new engines (Tesseract, TrOCR, etc.) in
- Image Preprocessing:
- Easily extend
utils/pdf_utils.pyto add more filters or preprocessing steps.
- Easily extend
- Output Formats:
- Add DOCX, HTML, or other formats by extending the output logic in the GUI and backend.
- Settings:
- Persistent via
QSettings. Add more options as needed.
- Persistent via
- Automated tests are provided in
tests/test_pdftextgenie.py. - Run all tests:
python -m pytest tests/
- Tests cover:
- Image preprocessing (binarize, denoise, contrast)
- Text cleanup (hyphens, paragraphs)
- OCR backend interface (mocked)
- Error handling for invalid files
- pdf2image errors:
- Ensure Poppler is installed and on your PATH (Windows).
- OCR accuracy issues:
- Try enabling preprocessing options or increasing DPI.
- GUI not launching:
- Check Python and PyQt5 installation.
- Other issues:
- Run tests and check error dialogs for details.
Contributions are welcome! To propose a feature, fix a bug, or add a new OCR backend:
- Fork the repo and create a new branch.
- Add your changes and tests.
- Open a pull request with a clear description.
Ideas for contribution:
- Add new OCR engines (Tesseract, TrOCR, LLaMA OCR)
- Multilingual support
- Export to DOCX/HTML
- Table/column structure preservation
- GUI enhancements and accessibility
- Author: @SK8-infi