AI Gateway
Every model.One API.
Interface your applications with LLMs and keep control of your costs, data and usage. One OpenAI-compatible API for 50+ providers, powered by the open-source Otoroshi LLM extension.
curl https://ai.my-company.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large",
"messages": [
{ "role": "user", "content": "Summarize this contract" }
]
}'Unified interface
One OpenAI-compatible API for every model. Simplify integrations, switch providers without touching your code.
50+ providers
OpenAI, Anthropic, Mistral, Gemini, Azure, Groq, DeepSeek, Scaleway, OVHcloud, Cohere, Hugging Face, Ollama…
Semantic cache
Speed up repeated queries, improve response times and reduce costs.
Load balancing & fallbacks
Distribute workloads across providers, retry and fall back automatically.
Quotas & budgets
Manage token quotas and budgets per consumer, team or API key, and keep costs in check.
Observability & reporting
Every LLM call is audited with its consumer, provider, usage, cost and ecological impact, and exportable anywhere.
One API for 50+ providers — cloud, sovereign or local
- OpenAI
- Anthropic
- Mistral AI
- Gemini
- Ollama
- Azure AI
- Groq
- DeepSeek
- Scaleway
- OVHcloud
- Cohere
- Hugging Face
- Cloudflare AI
- ElevenLabs
- + many more, local models too
FinOps for AI
Track and optimize
every LLM dollar.
Full visibility and control over the cost of every request. Monitor usage, generate cost reports per model, provider, API key or user, set budgets, and fine-tune your routing to reduce waste — cost tracking is enabled by default.
- Cost per call & per model
- Budgets per team
- Ecological impact
- Exportable audit events

Beyond chat completions
Everything you need to
run AI in production.
Guardrails
PII, secrets, prompt injection, toxicity and more, checked on requests and responses.
Prompts & contexts
Prompt templates, contexts and validation, managed centrally instead of hard-coded in every app.
MCP & agents
MCP connectors, virtual MCP servers, tool calling, workflows and AI agents, all governed by the gateway.
Multi-modal
Audio (text-to-speech, speech-to-text), images, embeddings and vector stores behind the same gateway.
Smart routing
Route by cost, latency or intent, load balance and fall back across providers and models.
Cost & CO₂ per call
Know the price and the ecological impact of every request, by model, provider or consumer.
AI Studio
The AI console
for your teams.
AI Studio gives your developers a self-service console on top of the gateway: workspaces, bring-your-own-key providers, API keys, a chat playground, routing, guardrails, logs and budgets — all stored as plain Otoroshi entities.
- Workspaces
- API keys & BYOK providers
- Chat playground
- Logs & activity

Run it your way
Managed, serverless
or self-hosted.
On Otoroshi Managed
The LLM extension is included in every managed plan, alongside every Otoroshi feature.
DiscoverOn Serverless
AI plugins are built in, no setup required. Start with the AI project starter.
DiscoverOn your Otoroshi
The Otoroshi LLM extension is open source (Apache 2.0). Install it on your own clusters.
DiscoverFAQ
Frequently asked questions
Still have a question? Talk to our team.
What is an AI gateway?
An AI gateway is similar to an API gateway, but designed for AI traffic. It manages, routes and secures calls to LLMs and other AI services, so you can integrate AI in your applications reliably and at scale.
Which AI providers can I use?
50+ providers through one OpenAI-compatible API, including OpenAI, Azure OpenAI, Anthropic, Mistral, Gemini, Groq, DeepSeek, Scaleway, OVHcloud AI Endpoints, Cohere, Hugging Face, Cloudflare AI, and local models with Ollama.
Where can I use the AI Gateway?
The AI features are built in on both Otoroshi Managed (the LLM extension is included in every plan) and Serverless. Since the LLM extension is open source, you can also run it on your own Otoroshi.
Can I route traffic to different LLMs?
Absolutely. Route requests based on rules like cost, performance or context, load balance across providers and fall back automatically when one fails.
How does semantic caching reduce AI costs?
Semantic caching identifies similar prompts and serves stored answers instead of calling the model again. It dramatically reduces the number of expensive model invocations and improves response times.
Can I set token quotas or budgets per API, team or model?
Yes. Define token quotas, request limits or budgets per route, API key, user or team to enforce spending limits across your organization.
Can I track LLM costs?
Yes. Cost tracking is enabled by default. You can monitor cost per request, generate spend reports per model or provider, and export them for billing or optimization.
Can I protect my prompts and data?
Yes. Guardrails detect and block PII, secrets leakage, prompt injection or toxic content on requests and responses, before anything reaches the model or your users.
Ready to build?Start in minutes.
Spin up a managed Otoroshi cluster or a serverless project for free, or ask us for a live demo of the whole platform.