Shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Plugins Filament, paquetes Laravel y starter kits open source mantenidos activamente. Todo MIT, con releases para v3, v4 y v5 cuando aplique.
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Instant, controllable, local pre-trained AI models in Rust
Self-hosted LLM gateway and tier-routing proxy for Claude Code, Cursor, and Codex. Routes across Ollama, AWS Bedrock, OpenRouter, Databricks, Azure OpenAI, llama.cpp, and LM Studio with prompt caching, MCP tools, and 60-80% cost savings.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
Instruct LLMs for flat and nested NER. Fine-tuning Llama and Mistral models for instruction named entity recognition. (Instruction NER)
Todos los paquetes tienen CI, tests Pest e issues abiertas marcadas con good first issue.
Nueva versión disponible.