Shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Plugins Filament, pacotes Laravel e starter kits open source mantidos ativamente. Tudo MIT, com releases para v3, v4 e v5 quando aplicável.
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Instant, controllable, local pre-trained AI models in Rust
Self-hosted LLM gateway and tier-routing proxy for Claude Code, Cursor, and Codex. Routes across Ollama, AWS Bedrock, OpenRouter, Databricks, Azure OpenAI, llama.cpp, and LM Studio with prompt caching, MCP tools, and 60-80% cost savings.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
Instruct LLMs for flat and nested NER. Fine-tuning Llama and Mistral models for instruction named entity recognition. (Instruction NER)
Todos os pacotes têm CI, testes Pest e issues abertas marcadas com good first issue.
Nova versão disponível.