Skip to content

AI Gateway

Every model.One API.

Interface your applications with LLMs and keep control of your costs, data and usage. One OpenAI-compatible API for 50+ providers, powered by the open-source Otoroshi LLM extension.

POST /v1/chat/completions
curl https://ai.my-company.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large",
    "messages": [
      { "role": "user", "content": "Summarize this contract" }
    ]
  }'
routedMistral · fallback OpenAI
guardrailsPII, prompt injection
budgetteam-legal · 62% used
impactcost & CO₂ per call

Unified interface

One OpenAI-compatible API for every model. Simplify integrations, switch providers without touching your code.

50+ providers

OpenAI, Anthropic, Mistral, Gemini, Azure, Groq, DeepSeek, Scaleway, OVHcloud, Cohere, Hugging Face, Ollama…

Semantic cache

Speed up repeated queries, improve response times and reduce costs.

Load balancing & fallbacks

Distribute workloads across providers, retry and fall back automatically.

Quotas & budgets

Manage token quotas and budgets per consumer, team or API key, and keep costs in check.

Observability & reporting

Every LLM call is audited with its consumer, provider, usage, cost and ecological impact, and exportable anywhere.

One API for 50+ providers — cloud, sovereign or local

  • OpenAI
  • Anthropic
  • Mistral AI
  • Gemini
  • Ollama
  • Azure AI
  • Groq
  • DeepSeek
  • Scaleway
  • OVHcloud
  • Cohere
  • Hugging Face
  • Cloudflare AI
  • ElevenLabs
  • + many more, local models too

FinOps for AI

Track and optimize
every LLM dollar.

Full visibility and control over the cost of every request. Monitor usage, generate cost reports per model, provider, API key or user, set budgets, and fine-tune your routing to reduce waste — cost tracking is enabled by default.

  • Cost per call & per model
  • Budgets per team
  • Ecological impact
  • Exportable audit events
otoroshi · AI Gateway — Costs
AI Gateway analytics dashboard showing spend, projected monthly spend and cost per call

Beyond chat completions

Everything you need to run AI in production.

Guardrails

PII, secrets, prompt injection, toxicity and more, checked on requests and responses.

Prompts & contexts

Prompt templates, contexts and validation, managed centrally instead of hard-coded in every app.

MCP & agents

MCP connectors, virtual MCP servers, tool calling, workflows and AI agents, all governed by the gateway.

Multi-modal

Audio (text-to-speech, speech-to-text), images, embeddings and vector stores behind the same gateway.

Smart routing

Route by cost, latency or intent, load balance and fall back across providers and models.

Cost & CO₂ per call

Know the price and the ecological impact of every request, by model, provider or consumer.

AI Studio

The AI console
for your teams.

AI Studio gives your developers a self-service console on top of the gateway: workspaces, bring-your-own-key providers, API keys, a chat playground, routing, guardrails, logs and budgets — all stored as plain Otoroshi entities.

  • Workspaces
  • API keys & BYOK providers
  • Chat playground
  • Logs & activity
ai-studio · Routing
AI Studio routing screen

Run it your way

Managed, serverless or self-hosted.

FAQ

Frequently asked questions

Still have a question? Talk to our team.

What is an AI gateway?

An AI gateway is similar to an API gateway, but designed for AI traffic. It manages, routes and secures calls to LLMs and other AI services, so you can integrate AI in your applications reliably and at scale.

Which AI providers can I use?

50+ providers through one OpenAI-compatible API, including OpenAI, Azure OpenAI, Anthropic, Mistral, Gemini, Groq, DeepSeek, Scaleway, OVHcloud AI Endpoints, Cohere, Hugging Face, Cloudflare AI, and local models with Ollama.

Where can I use the AI Gateway?

The AI features are built in on both Otoroshi Managed (the LLM extension is included in every plan) and Serverless. Since the LLM extension is open source, you can also run it on your own Otoroshi.

Can I route traffic to different LLMs?

Absolutely. Route requests based on rules like cost, performance or context, load balance across providers and fall back automatically when one fails.

How does semantic caching reduce AI costs?

Semantic caching identifies similar prompts and serves stored answers instead of calling the model again. It dramatically reduces the number of expensive model invocations and improves response times.

Can I set token quotas or budgets per API, team or model?

Yes. Define token quotas, request limits or budgets per route, API key, user or team to enforce spending limits across your organization.

Can I track LLM costs?

Yes. Cost tracking is enabled by default. You can monitor cost per request, generate spend reports per model or provider, and export them for billing or optimization.

Can I protect my prompts and data?

Yes. Guardrails detect and block PII, secrets leakage, prompt injection or toxic content on requests and responses, before anything reaches the model or your users.

Ready to build?Start in minutes.

Spin up a managed Otoroshi cluster or a serverless project for free, or ask us for a live demo of the whole platform.