Vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Plugins Filament, paquetes Laravel y starter kits open source mantenidos activamente. Todo MIT, con releases para v3, v4 y v5 cuando aplique.
A high-throughput and memory-efficient inference and serving engine for LLMs
The open-source AI voice studio. Clone, dictate, create.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.
Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU host cache, cross-layer expert prefetch, and the first MTP speculative decoding for GLM-5.2.
Fine-tune Piper (VITS) text-to-speech voices on any CUDA GPU or NVIDIA DGX Spark — reproducible setup, checkpoint-resume handoff, and the piper1-gpl build gotchas solved. Bring your own dataset.
Todos los paquetes tienen CI, tests Pest e issues abiertas marcadas con good first issue.
Nueva versión disponible.