PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Filament plugins, Laravel packages and open source starter kits actively maintained. All MIT, with releases for v3, v4 and v5 where applicable.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
A community-supported supercharged document management system: scan, index and archive all your documents
OpenDataLoader PDF - Monorepo workspace
A fast, helpful, and open-source document parser
📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
One delightful Ruby framework for every major AI provider. Build AI agents, chatbots, RAG apps, and multimodal workflows in beautiful, expressive code.
An efficient OCR engine for receipt image processing.
C++ implementations of PP-OCRv3/4/5/6 using ncnn for inference.
Laravel OCR & Document Data Extractor - A powerful OCR and document parsing engine for Laravel
All packages have CI, Pest tests and open issues tagged good first issue.
New version available.