Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Plugins Filament, packages Laravel et starter kits open source activement maintenus. Tous en MIT, avec des releases pour v3, v4 et v5 selon les cas.
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of OpenAI's Whisper model in C/C++
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
C++ implementations of PP-OCRv3/4/5/6 using ncnn for inference.
Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI + systemd.
Tous les packages ont une CI, des tests Pest et des issues ouvertes taguées good first issue.
Nouvelle version disponible.