If your app talks to more than one language model, you eventually need a middleman: one endpoint, one API key, and one place to watch costs and failures. That middleman is an AI API gateway. Since Stripe's announced acquisition of OpenRouter in August 2026, a lot of developers have started mapping exit routes — and the good news is that routing models is now a solved, well-priced problem with genuinely free options on both the self-hosted and managed sides.
What an AI gateway actually does
An AI gateway is a proxy between your code and LLM providers. You point your OpenAI-compatible client at the gateway's base_url instead of a provider's, and the gateway handles routing to the right model, automatic failover when one provider is down, response caching, rate limiting, and per-request cost tracking. The whole pitch is drop-in compatibility: change one line, keep the rest of your code.
The practical difference between gateways comes down to three questions: who holds your API keys, what markup you pay on tokens, and how much operational work you accept.
The 7 free gateways worth using
1. LiteLLM — the free self-hosted standard
LiteLLM is the default open-source gateway: a Python proxy under the MIT license that speaks OpenAI-compatible endpoints to more than 100 providers. You run it yourself, your keys never leave your infrastructure, and the gateway itself charges zero markup — you only pay the providers' list prices for tokens.
The trade-off is operational: you install it, you scale it, you patch it. For side projects and teams that already run servers, that is usually worth it. It also doubles as a spend-control layer, letting you set per-key budgets and track usage per user or project.
2. Vercel AI Gateway — zero markup, managed
Vercel's gateway is the cleanest hosted option for developers who want zero ops. It routes to hundreds of models and charges no markup on tokens — you pay the providers' list prices, plus whatever card processing your own payment setup incurs. It works especially well with the Vercel AI SDK, so Next.js apps get first-class integration.
The main constraint is ecosystem: it is at its best when you already deploy on Vercel, and managed gateways always see your prompts, so teams with strict data-handling requirements should check the terms before routing sensitive traffic through it.
3. Cloudflare AI Gateway — free control plane over your own keys
Cloudflare's AI Gateway is a free observability and routing layer that sits in front of your existing provider keys. Core features — request logging, caching, rate limiting, and fallback routing — are free on all plans, and there is no markup when you bring your own provider keys.
Its strength is placement: it lives on Cloudflare's edge, which makes it a natural pick if your app already runs on Cloudflare infrastructure. It also offers DLP and PII scanning features, useful if you want basic prompt hygiene before traffic reaches a model provider.
4. Portkey — free tier with guardrails built in
Portkey's cloud plan starts with a free tier of 10,000 logs per month, stepping up to a $49/month Production plan. Beyond routing, Portkey is aimed at teams that need governance: guardrails, prompt templates, and compliance tooling are part of the pitch. There is also an open-source gateway edition you can self-host.
Choose Portkey when your team needs more than routing — RBAC, SSO, and audit trails for production LLM traffic. For a single developer just routing cheap models, the free tier still works fine as an observability layer.
5. Helicone — observability-first, free to start
Helicone started as an observability proxy for LLM calls and has grown into a full AI gateway. Its free tier covers a generous volume of requests (its paid plan starts at $79/month), and it is open source. If your main worry is understanding where your money goes — per-model cost, latency, and error rates — Helicone's dashboards are the most developer-friendly of the bunch.
Pair it with any provider: because it is a proxy, you can route through Helicone and still keep your existing provider relationships and volume discounts.
6. Bifrost — self-hosted speed in Go
If LiteLLM is the feature-rich Python option, Bifrost is the performance play: a self-hosted gateway written in Go, built for high throughput with a small memory footprint. Like LiteLLM it charges no markup and keeps your keys on your own infrastructure.
Reach for Bifrost when your gateway is in a hot path — high-concurrency agents, batch pipelines — and you want LiteLLM's job done with lower latency and a thinner advisory record. The trade-off is a smaller feature surface than the Python ecosystem.
7. Kong AI Gateway — the enterprise self-hosted option
Kong's AI Gateway builds LLM routing on top of the mature Kong API gateway, with a 30-day trial for the commercial tier and an open-source core. This is the pick when you are already a Kong shop: you get AI traffic management alongside your existing APIs, with enterprise features like PII sanitization and RBAC available on the paid tiers.
It is heavier than LiteLLM or Bifrost for a solo developer, but for Kubernetes-native teams with governance requirements, it slots into infrastructure you already understand.
Quick comparison
| Gateway | Type | Markup | Free tier | Best for |
|---|---|---|---|---|
| LiteLLM | Self-hosted (Python) | $0 | Free forever (MIT) | Self-hosters wanting full control |
| Vercel AI Gateway | Hosted | 0% | Usage-based, no markup | Vercel/Next.js developers |
| Cloudflare AI Gateway | Hosted | 0% with own keys | Free core features | Apps already on Cloudflare |
| Portkey | Hosted + OSS | Free tier, then $49/mo | 10K logs/mo | Teams needing guardrails |
| Helicone | Hosted + OSS | Free tier, then $79/mo | Generous free requests | Observability-first teams |
| Bifrost | Self-hosted (Go) | $0 | Free forever | High-throughput self-hosting |
| Kong AI Gateway | Self-hosted | $0 core | 30-day trial | Enterprise Kong/K8s shops |
How to switch: it is one line
The gateway pattern is deliberately boring. Here is the standard drop-in swap with the OpenAI SDK:
from openai import OpenAI
# Point at the gateway, keep everything else
client = OpenAI(
base_url="https://ai-gateway.vercel.sh/v1",
api_key="your-gateway-key",
)
resp = client.chat.completions.create(
model="anthropic/claude",
messages=[{"role": "user", "content": "Hello!"}],
)For a self-hosted LiteLLM proxy, the same code points at http://localhost:4000. That uniformity is the point: once your code talks to a gateway, moving between providers — or between gateways — stops being a rewrite.
Where OpenRouter fits in 2026
OpenRouter remains the zero-ops hosted option with 400+ models and roughly 5.5% markup on credit top-ups. But it is no longer the only sensible default. Two of the gateways above charge no markup at all, and one of them is free software you can run on a $5 VPS. If the Stripe acquisition makes you uneasy about single-vendor lock-in, mapping a gateway exit now is cheap insurance: the switch itself is one line of code.
Picking the right one
- Zero ops, zero markup: Vercel AI Gateway or Cloudflare AI Gateway.
- Your keys, your infra, $0 forever: LiteLLM (broadest features) or Bifrost (fastest).
- You need governance and audit trails: Portkey or Kong AI Gateway.
- You mainly want to see where money goes: Helicone.
- EU data residency matters: look at EU-first routers such as Requesty or Eden AI alongside these.

