Llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Aktiv gepflegte Filament-Plugins, Laravel-Pakete und Open-Source-Starter-Kits. Alle MIT, mit Releases für v3, v4 und v5, sofern zutreffend.
Hundreds of models & providers. One command to find what runs on your hardware.
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU host cache, cross-layer expert prefetch, and the first MTP speculative decoding for GLM-5.2.
Alle Pakete haben CI, Pest-Tests und offene Issues mit dem Tag good first issue.
Neue Version verfügbar.