Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
Point your existing SDK at one URL and reach every LLM vendor — with real failover, not a try/except. One static Rust binary.
OpenAI/Anthropic-compatible AI router that keeps apps online with provider failover, key rotation, caching, analytics, and local-model fallback.