Llmfit
Hundreds of models & providers. One command to find what runs on your hardware.
Plugins Filament, packages Laravel et starter kits open source activement maintenus. Tous en MIT, avec des releases pour v3, v4 et v5 selon les cas.
Hundreds of models & providers. One command to find what runs on your hardware.
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU host cache, cross-layer expert prefetch, and the first MTP speculative decoding for GLM-5.2.
Tous les packages ont une CI, des tests Pest et des issues ouvertes taguées good first issue.
Nouvelle version disponible.