PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Plugins Filament, packages Laravel et starter kits open source activement maintenus. Tous en MIT, avec des releases pour v3, v4 et v5 selon les cas.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OpenDataLoader PDF - Monorepo workspace
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
A fast, helpful, and open-source document parser
PDFX — a backwards-compatible PDF extension for multi-document files, with a minimal viewer
Tous les packages ont une CI, des tests Pest et des issues ouvertes taguées good first issue.
Nouvelle version disponible.