Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Plugins Filament, packages Laravel et starter kits open source activement maintenus. Tous en MIT, avec des releases pour v3, v4 et v5 selon les cas.
A high-throughput and memory-efficient inference and serving engine for LLMs
The open-source AI voice studio. Clone, dictate, create.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU host cache, cross-layer expert prefetch, and the first MTP speculative decoding for GLM-5.2.
Fine-tune Piper (VITS) text-to-speech voices on any CUDA GPU or NVIDIA DGX Spark — reproducible setup, checkpoint-resume handoff, and the piper1-gpl build gotchas solved. Bring your own dataset.
Tous les packages ont une CI, des tests Pest et des issues ouvertes taguées good first issue.
Nouvelle version disponible.