# OneVir by OneQuill > Self-hosted AI gateway and LLM inference engine in one native server process. Runs where your data lives. ## Product - [OneVir](https://onevir.onequill.dev/): product overview and demonstrations. - [Capabilities](https://onevir.onequill.dev/#capabilities) - [Architecture](https://onevir.onequill.dev/#architecture) - [Company](https://onequill.dev/about) OneVir serves GGUF models on CPU or GPU using llama.cpp, Vulkan or CUDA. It routes to 26 types of cloud provider only through configured routes. Applications use OpenAI-compatible or Anthropic-compatible APIs. Capabilities include model placement, aliases and ordered fallbacks, request policy, guard models, PII redaction, tool approvals, signed kill switch, emergency stop, per-key limits, provider budgets, Prometheus metrics and opt-in OpenTelemetry export. Speech support uses Whisper, Piper and Kokoro. Telemetry excludes prompts, completions, tool payloads and keys. Validated on Windows 11. Linux and macOS deployment paths exist but are not certified. Central fleet management, enterprise SSO, full administrative RBAC, multi-tenancy and the OpenAI Responses API are roadmap items. ## Licensing and support - [Licensing](https://onequill.dev/licensing): Business Source License 1.1. Commercial production requires a commercial licence; non-commercial use and non-production evaluation are free under the licence. - [Binding licence text](https://onequill.dev/LICENSE) - [Support](https://onequill.dev/support) - [Privacy](https://onequill.dev/privacy) Sales and walkthroughs: sales@onequill.dev. Support: support@onequill.dev. Last updated: 2026-10-02.