A self-hosted PDF OCR API that converts scanned documents to markdown. Powered by PaddleOCR-VL, runs on GPU via Docker.
-
Updated
Jul 12, 2026 - Python
A self-hosted PDF OCR API that converts scanned documents to markdown. Powered by PaddleOCR-VL, runs on GPU via Docker.
Pdf utilities for text extraction in digital and convert scanned pdf into canvas.
为PDF格式的电子书生成目录。通过多模态AI实现,自动生成有层次的目录书签。专为扫描版/纯图片PDF设计。 Generate a table of contents for PDF ebooks via multimodal AI. Auto-produces hierarchical bookmarks. Designed for scanned / image-only PDFs.
Convert Word/Excel/PowerPoint to PDF with custom watermarks and non-editable scanned output. 100% offline & private - documents never leave your PC. Free, no account.
Medical OCR refinement pipeline for Codex / 医学文本OCR Skill
Outil OCR permettant d’extraire et de structurer du texte à partir d’images et de PDF scannés (export en .docx et .txt) — prise en charge du français et de l’anglais
Lightweight bash script to convert scanned PDFs into searchable, copyable PDFs using Tesseract OCR with parallel processing.
Turn the scanned PDFs you own into clean, reflowable EPUB3 — studio-grade Apple Vision OCR that runs on your Mac, not someone's cloud. 100% offline, free, MIT.
Add a description, image, and links to the scanned-pdf topic page so that developers can more easily learn about it.
To associate your repository with the scanned-pdf topic, visit your repo's landing page and select "manage topics."