Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Aktiv gepflegte Filament-Plugins, Laravel-Pakete und Open-Source-Starter-Kits. Alle MIT, mit Releases für v3, v4 und v5, sofern zutreffend.
A high-throughput and memory-efficient inference and serving engine for LLMs
Packs ECMAScript/CommonJs/AMD modules for the browser. Allows you to split your codebase into multiple bundles, which can be loaded on demand. Supports loaders to preprocess files, i.e. json, jsx, es7, css, less, ... and your custom stuff.
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Alle Pakete haben CI, Pest-Tests und offene Issues mit dem Tag good first issue.
Neue Version verfügbar.