Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Filament plugins, Laravel packages and open source starter kits actively maintained. All MIT, with releases for v3, v4 and v5 where applicable.
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of OpenAI's Whisper model in C/C++
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
C++ implementations of PP-OCRv3/4/5/6 using ncnn for inference.
Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI + systemd.
All packages have CI, Pest tests and open issues tagged good first issue.
New version available.