For a team already juggling DeepSeek, Kimi, Qwen, or other providers, a shared routing layer can be reasonable if it removes duplicated work and the fallback behavior is tested. GPTProto is one implementation of this approach.
Still paying list price and getting throttled?Pay 10–30% less.
Point Claude Code and Codex CLI at one endpoint. Same models, same protocol, same SDK — 10–30% below each vendor’s list price. Pay per token, no minimum, no subscription.
Set up GPTProto for Claude Code.
1. In ~/.claude/settings.json, under "env" set:
ANTHROPIC_BASE_URL = "https://gptproto.com"
ANTHROPIC_API_KEY = "sk-gptproto-<your key>"
Keep it exactly like that - no /v1, no trailing slash. Claude Code appends /v1/messages itself, so a /v1 here becom...The Same Models.A Smaller Bill.
Automation built on agent frameworks — coding, office ops, content distribution and scheduled jobs running unattended.
Tool callingOpenAI-compatibleHigh throughputReliable routing| Model | Price · GPTProto | vs Official | vs OpenRouter | Context | Modality | Stability | Action |
|---|---|---|---|---|---|---|---|
| $0.12 / $0.30$0.03 cache read · per 1M tokens | −40% | −43% | — | → | Try | ||
| $0.16 / $0.65per 1M tokens | −40% | −43% | — | → | Try | ||
| $0.16 / $0.96$0.02 cache read · per 1M tokens | −20% | −24% | 1.05M | → | Try | ||
| $0.30 / $1.20$0.01 cache read · per 1M tokens | — | −5% | 1.05M | → | Try | ||
| $1.80 / $9.00per 1M tokens | −40% | −43% | — | → | Try | ||
| $1.26 / $3.96$0.23 cache read · per 1M tokens | −10% | −15% | 1.05M | → | Try | ||
| $2.70 / $13.50$0.27 cache read · per 1M tokens | −10% | −15% | 1M | → | Try | ||
| $0.60 / $3.60$0.06 cache read · per 1M tokens | −20% | −24% | 400K | → | Try | ||
| $6.40per image | −20% | −24% | — | → | Try | ||
| $0.28 / $1.12$0.07 cache read · per 1M tokens | −30% | −34% | 1.05M | → | Try | ||
| $4.50 / $22.50$0.45 cache read · per 1M tokens | −10% | −15% | 1M | → | Try | ||
| $0.18 / $1.50$0.02 cache read · per 1M tokens | −40% | −43% | 1.05M | → | Try | ||
| $0.10 / $0.42$0.05 cache read · per 1M tokens | −30% | −34% | 128K | → | Try | ||
| $0.48 / $0.96$0.10 cache read · per 1M tokens | −20% | −24% | 1.05M | → | Try | ||
| $1.80 / $9.00$0.18 cache read · per 1M tokens | −10% | −15% | 1M | → | Try | ||
| $0.75 / $1.50$0.12 cache read · per 1M tokens | −40% | −43% | 1M | → | Try | ||
| $1.20 / $3.60$0.30 cache read · per 1M tokens | −40% | −43% | 500K | → | Try | ||
| $0.30 / $1.20$0.01 cache read · per 1M tokens | — | −5% | — | → | Try | ||
| $0.79 / $2.38$0.04 cache read · per 1M tokens | −5% | −10% | 1.05M | → | Try | ||
| $8.00 / $40.00$0.80 cache read · per 1M tokens | −20% | −24% | 1.05M | → | Try |
Estimate the cost.
See the savings.
Choose a workflow and usage level to see an example monthly cost on GPTProto, alongside the same usage at comparable provider rates.
Batch work and long-context runs burn credits faster — the recommendation adjusts automatically as you switch workload or usage level.
Same CLI, same protocol.Two lines change.
Two environment variables for Claude Code, three lines of config.toml for Codex CLI. Nothing else in your project moves.
2 lines changed · 0 code rewrites · 200+ models on one key · model names stay the same.
One host. One key.
Failover included.
Both CLIs point at the same origin. One balance covers every client, and a dropped channel does not stop the session.
Both protocols, one host.
Anthropic-native and OpenAI-compatible routes are served from one origin, so both CLIs point at the same host. No gateway, no adapter, no extra SDK.
One key, one balance, one bill.
Every CLI, cron job and agent draws on the same balance, at the same discounted rates. Top up once — bonus credits spend at those rates too.
See top-up tiersAuto-failover on redundant channels.
Major models run on redundant channels. If one hits its limit or goes down, traffic shifts to a backup — the ceiling is that channel’s, not the platform’s.
Your Usage.Your Potential Savings.
Choose a model and enter your monthly usage to compare estimated costs on GPTProto, the official provider, and OpenRouter.
Don't take our word for it.Take theirs.
Real posts from real, public accounts — nothing here is invented.

I've been testing GPTProto recently, and it's honestly made my creative workflow much simpler. Instead of paying for multiple subscriptions, I can access several leading AI models from one place.

I turned this single prompt into a cinematic fantasy video using GPTProto. "Continuous 15-second cinematic shot, 4K resolution, hyper-realistic dark fantasy photorealism…"

I challenged myself to create a cinematic AI cooking short in just 15 seconds. I used GPTProto to bring the entire workflow together, from image generation to video, all in one place.

Made with Seedance 2.0 + GPT Image 2 on GPTProto. A Pixar-style commercial with the perfect glow.
What worked for me was pointing the cloud connections at GPTProto so the frontend only sees one endpoint and I just change the model name to swap. I am not rebuilding a connection from scratch every time.
I route the calls through GPTProto so a fallback is a config switch instead of a weekend rewrite when a model disappears. It turns "my default model just got export controlled" from an incident into a config change.

I used to switch between different AI tools just to compare results. Now I just use GPTProto. GPT-5, Claude, Gemini, Kimi, and more — all in one workspace.
I route the calls through GPTProto so the model id and latency land in one place regardless of which provider is behind it. The win is having the log schema consistent across providers.

I created this 15-second cinematic product video with GPTProto using Seedance 2.0, and I was really impressed by how smooth the workflow was.

GPTProto routes the character prompt and shot list to a top model on one API key: 20 minutes … voices it and burns in captions from the same key: 20 minutes.
Text, images, and voice are configured separately. You can keep OpenRouter for text and use GPTProto for visuals.
The call layer underneath the router is GPTProto, so swapping a model does not require provisioning a new provider integration. Changing models becomes cheap enough that the question stops being "should we change."
Pay Per Token.
No Subscription. No Minimum.
You only pay for what you use. One balance covers Claude Code, Codex CLI and every other client you run — top-ups of $20+ earn bonus credits that spend at the same discounted rates.
Enough to point one CLI at the endpoint and run a real session. Top up $20+ anytime to unlock bonus credits.
The sweet spot for a developer running Claude Code and Codex CLI every day without watching the meter.
Built for a fleet of CLIs and agent runners. Bonus credits stack on top of already-discounted model pricing, so the whole balance prices 10–30% under list.
Before You Edit the Config.Straight Answers.
Short answers. No sales talk.
01Which CLIs can I point at GPTProto?
Any client that speaks one of the two protocols OpenAI and Anthropic already defined. That covers Claude Code, Codex CLI, OpenCode, Cline, Kilo Code and most OpenAI-SDK agent frameworks out there. You choose the protocol by choosing the base URL — nothing gets installed in between.
02Why does Claude Code take a different base URL than Codex CLI?
Because the two clients append different paths. Claude Code appends /v1/messages to whatever you set, so its base URL must be https://gptproto.com — no /v1, no trailing slash. Add /v1 and the request goes to /v1/v1/messages, which returns 404. Codex CLI is the mirror image: it expects the OpenAI-compatible route, which only exists under /v1/chat/completions, so its base_url is https://gptproto.com/v1.
03Do I have to change any code?
No. Both endpoints are protocol-compatible, so your SDK, streaming, tool calls, function schemas and the CLI binaries themselves are untouched. What changes is a base URL and an API key — environment variables or one config file. Most people are running in under ten minutes.
04Do the model IDs change?
The IDs match the underlying models, so the ones you already use — claude-opus-5, claude-sonnet-5, gpt-5.2-codex, glm-5.2 — keep working. GET /v1/models on the same host returns the full list if you want to check a specific one before switching.
05Why is it 10–30% cheaper than official?
No seat fee and no fixed plan — you pay per token against upstream capacity we buy in bulk, and the discount is passed through as a flat multiplier per model. Claude models are ~10% under list, the Codex family ~30%, and everything in between. The rate is the rate: it does not shrink as you use more.
06Do bonus credits stack with the model discount?
Yes. Bonus credits from a top-up spend at the same discounted model prices, so the 10–30% applies across your whole balance — including the credits you did not pay for. A $500 top-up lands as $600 of balance that buys $600 at discount rates.
07What happens when an upstream provider rate-limits or goes down?
Requests route through redundant upstream channels with automatic failover. If a channel 429s or drops, traffic shifts to a backup rather than failing your request. You can still pin your own model fallback in the request if you want deterministic routing.
08Are there rate limits on GPTProto itself?
There is no per-minute ceiling you have to plan around the way you do on a single vendor’s plan — throughput comes from the pooled channels behind each model. Fair-use limits apply per account; if a workload needs sustained very high concurrency, contact us from your registered email and we will size it.
09Can I switch back if it doesn’t work out?
Yes, trivially. Nothing is installed and nothing was rewritten, so reverting is changing the same two values back. Your balance stays in the account and there is no subscription to cancel or term to sit out.
Read the Docs.Skip the Guesswork.
Exact base URLs, model IDs and per-endpoint examples for both protocols.
Your next refactorshouldn't run at list price.
Create images and videos online, or bring text, image, and video models into your app with one API key. Explore discounted rates on selected models.
- ✓Both protocols, one host
- ✓200+ models on one key
- ✓10–30% below official
- ✓High uptime, auto-failover