Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Aktiv gepflegte Filament-Plugins, Laravel-Pakete und Open-Source-Starter-Kits. Alle MIT, mit Releases für v3, v4 und v5, sofern zutreffend.
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of OpenAI's Whisper model in C/C++
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
C++ implementations of PP-OCRv3/4/5/6 using ncnn for inference.
Open-source self-hosted home AI inference platform for AMD Strix Halo — multi-backend slots, OpenAI-compatible gateway, Vue 3 + FastAPI + systemd.
Alle Pakete haben CI, Pest-Tests und offene Issues mit dem Tag good first issue.
Neue Version verfügbar.