Engineering · 7 min read ·
LLM Gateway vs OpenRouter: When to Self-Host Your AI Gateway
One OpenRouter key shared by every script and teammate is convenient until it isn't. Here's when to put your own gateway in front.

OpenRouter gives you one OpenAI-compatible API for hundreds of models from many providers, with one bill. For a single developer it is hard to beat. Problems start when ten scripts, three agents and five teammates share the same key: nobody knows who spent what, one runaway loop can exhaust the budget, and rotating the key breaks everything at once.
That is the job of an LLM gateway.
OpenRouter vs an LLM gateway
| OpenRouter | Self-hosted LLM gateway | |
|---|---|---|
| Main job | Access to many models through one API | Control over who uses models, how much and how |
| Keys | Your account's keys | Its own per-consumer keys; the upstream key stays hidden |
| Limits | Account-level | Per team, app or agent |
| Logs | Account activity | Your logs, in your infrastructure |
| Routing | Provider routing and fallbacks | Your own rules (e.g. "auto" → default model) |
They are complementary. A common setup puts a self-hosted gateway in front of OpenRouter: OpenRouter handles the model catalogue, the gateway handles your organisation.
When a self-hosted gateway pays off
- More than one team, product or agent shares model access.
- You need spend attribution per client or project.
- Upstream keys must never reach laptops, CI logs or client apps.
- You want to switch default models without redeploying every app.
- Compliance requires request logs you control.
What a gateway should do
- Issue its own keys, one per consumer, that can be revoked individually.
- Keep the upstream key encrypted at rest and never return it.
- Rate-limit per key (requests per minute, tokens or spend per day).
- Route models — aliases like
autoorcheapthat map to real models you can change centrally. - Log usage per key: model, tokens, latency, status.
- Stay OpenAI-compatible, so existing SDKs work by changing only the base URL and key.
A reference implementation
Our open-source openrouter-api-gateway is a FastAPI service that turns one OpenRouter account into a multi-tenant proxy. Admins create sk-gw-… keys per consumer, set per-key rate limits and routing rules, and watch live request logs. Requests to /v1/* are forwarded upstream with the stored OpenRouter key, which is encrypted with Fernet; admin passwords use bcrypt and sessions are HMAC-signed.
Using it from an app is just a base URL change:
from openai import OpenAI
client = OpenAI(base_url="https://ai-gateway.example.com/v1", api_key="sk-gw-…")
client.chat.completions.create(model="auto", messages=[{"role": "user", "content": "Hi"}])
Alternatives
If you'd rather not run it yourself, LiteLLM's proxy, Portkey and Cloudflare AI Gateway cover similar ground with different trade-offs in hosting, pricing and features. The deciding questions are the same: where do keys and logs live, and how granular are limits?
Security notes
- Serve the gateway over HTTPS only and put the admin console behind SSO or an IP allowlist.
- Never log full prompts containing personal data unless you need them and are allowed to keep them.
- Alert on spend spikes per key — it is the fastest way to catch a looping agent.
Running AI across several teams or clients? We design and host AI infrastructure with access control and cost reporting built in.
Frequently asked questions
Is OpenRouter an LLM gateway?
It is a hosted gateway to many model providers. A self-hosted gateway adds organisation-level control — per-team keys, limits and logs — and can sit in front of OpenRouter.
Do I need to change my code to use a gateway?
Usually not. OpenAI-compatible gateways only need a new base URL and a gateway key.
How do I stop one agent from using the whole AI budget?
Give each agent its own gateway key with request and spend limits, and alert when usage jumps above its normal range.

