Your coding tool wants an API endpoint. Your providers each want their own auth, their own format, their own quota rules. 9Router sits in the middle: a self-hosted router that gives every tool one OpenAI-compatible /v1 endpoint and decides, per request, which provider should answer.
The problem it solves
If you code with AI agents, you already know the friction:
- Subscription quota expires unused. You pay Claude Code or Copilot monthly and lose whatever you did not burn.
- Rate limits stop you mid-task. One provider throttles, and the agent just dies.
- Tool outputs eat tokens.
git diff,grep,ls,treeoutputs routinely make up a third of the prompt budget — you pay for tokens that carry no reasoning value. - Each provider speaks a different dialect. OpenAI, Anthropic, Gemini, Cursor, Kiro, Vertex — every client expects its own wire format.
9Router is the decolua/9router project packaged for self-hosting: a single service that fronts all of it.
What it actually does
Point any AI coding tool at https://your-pod.generated-domain.com/v1 and it behaves like an OpenAI-compatible API. Behind that endpoint, 9Router runs the routing layer:
- Smart three-tier fallback. Route Subscription providers first (use the quota you already paid for), then cheap APIs, then free tiers — automatically, per request. One tier exhausts, the next answers. Zero downtime.
- Format translation. OpenAI ↔ Claude ↔ Gemini ↔ Cursor ↔ Kiro ↔ Vertex. Your tool sends one format; each provider gets its native one.
- Quota tracking. Live token counts plus a reset countdown per provider — you see exactly how much subscription value is left before it expires.
- Multi-account round-robin. Several accounts on one provider share the load.
- Auto token refresh. OAuth sessions renew themselves; no manual re-login.
- Custom combos. Name a model chain —
my-coding-stack= Claude → GLM → free fallback — and call it like a single model. - Usage analytics and request logging. See tokens, cost, and trends; debug mode dumps full request/response for troubleshooting.
The token savers
This is the part that justifies the name on the tin. Three optional savers run before the request ever reaches a provider:
- RTK Token Saver compresses tool outputs —
git-diff,grep,find,ls,tree, log dumps — before they hit the model. Auto-detected per output, lossless, silently skipped if compression would not help. Typical saving: 20–40% of input tokens. - Caveman mode injects a terse-reply prompt so the model answers in compressed technical shorthand. Up to 65% off output tokens.
- Ponytail biases the model toward minimal, YAGNI-first code — fewer output tokens and less unrequested scaffolding.
Savers are per-request controllable: send X-9Router-Token-Saver: off to bypass them for one call.
Which tools it serves
Anything that accepts an OpenAI-compatible endpoint: Claude Code, Codex, Cursor, Cline, Copilot, OpenCode, Kilo Code, Continue, Roo, Droid, DeepSeek TUI, Qwen Code, Devin CLI — set the base URL and API key and it works. The dashboard issues keys, tracks usage per key, and lets you name combos the tools can call by alias.
Why it wants a Pod
A router only helps if it is always reachable. Run it on a laptop and it dies with the lid; run it on a random VPS and you are back to Docker, TLS, and reverse-proxy setup — the exact chores a router exists to abstract.
On Cadesia, 9Router is a Pod: pick it in the App Library, name the Pod, create it. You get the dashboard and the /v1 endpoint on the Pod’s own generated HTTPS Address, persistent storage for provider config, combo definitions, and usage history, and the managed runtime keeping it up — for $3.99/mo on the Starter plan.
From there, every tool on every machine you own points at one URL:
Endpoint: https://your-pod.generated-domain.com/v1
API key: <from the 9Router dashboard>
Model: kr/claude-sonnet-4.5 # or any combo alias
One Address, every provider, and the token savers running on all of it.