Every price below measured

Every price we charge, next to the price you're paying now.

sagerouter is a metered, OpenAI-compatible API for open-weight models — GLM, Qwen, Kimi and friends. We publish our rate and the prevailing market rate side by side on every model we list, including the ones where the win is small. Then we let you set a hard cap so the bill can never surprise you.

No card. Prepaid credits only, so there is no invoice to be shocked by.

one line changes
# the request you already send OpenAI
curl https://sagerouter.ai/v1/chat/completions \
  -H "Authorization: Bearer sk-sr-..." \
  -d '{"model":"glm-5.2","messages":[...]}'

# what comes back — metered, itemised
"usage": {
  "prompt_tokens":     1842,
  "completion_tokens": 214,
  "cost_usd":          
}

Same request you send OpenAI. Different base_url. The cost above is that call priced at glm-5.2's rate in the table below — computed on this page from the same numbers, not typed in.

Serving today

The whole price list. Nothing withheld.

Per 1M tokens, USD. "Market" is the prevailing published rate for the same model on the major aggregators. We are cheaper on almost every line — sometimes by a lot, sometimes by a little. We show you which is which, because a vendor who hides the small numbers is hiding something.

ModelMarket in / outsagerouter in / outYou save

Prices are prepaid and metered per token. No minimum, no monthly platform fee, no seat licence. Models we cannot currently serve are not listed — we would rather show you a shorter list than sell you a name that returns an error. The roster grows as we validate capacity, and every addition ships with its market comparison on day one.

What would you actually pay?

Two numbers to move. The answer updates as you drag. This is the same arithmetic that runs on your real invoice — there is no second, worse price waiting after signup.

60M
25%
sagerouter $0
Same model, market rate $0
cheaper than an Opus-class frontier model at this volume ($0/mo)

The gap between the two bars is what sagerouter saves you. The gap to the frontier number is what open-weight models save you — that one is not our doing, and we are not going to pretend it is.

Your bill
$0
per month, metered
vs market
$0
under market
Hard cap
You set it
we stop at the number

The bill cannot surprise you

The failure everyone has a story about is a runaway agent loop burning a month of budget over a weekend. Set a hard monthly cap and we stop serving at the cap — not a warning email, not an overage line, a refusal with the reason in it.

Request after the cap is reached — HTTP 402
# your agent loop, 3am Saturday { "error": { "message": "This account has reached its monthly spend limit. Raise the limit to continue; the limit resets at the start of next month.", "type": "monthly_limit_reached", "request_id": "req_..." } }

Credits are prepaid, so the cap is a second floor under a floor: we cannot bill you for money you have not already put in. The error names the limit rather than the balance, so nobody gets sent to add funds that would not have helped.

Where the discount actually comes from

Cheap inference with no explanation should worry you. Here is ours, in plain language.

1. Open weights, not frontier rent

Most of the saving isn't ours — it's the model class. GLM-, Qwen- and Kimi-class weights cost a fraction of a closed frontier model to serve. Any router will give you that. We say so instead of billing you for it as if we invented it.

2. We buy on discounted supply

This is the part that is actually ours. We source each model on the cheapest channel we can verify by billing it — one real metered call per model, then read the coefficient off the supplier's own billing log, rather than trusting a rate card. That is where the rest of the gap comes from.

3. One multiplier, the same for everyone

Your price is our measured cost times a fixed multiplier. It does not vary by model, by customer, or by how much you spend — there is no volume tier you are failing to qualify for. When our cost falls, the table above falls with it and you renegotiate nothing.

Your data, and what we do with it

You're routing client work through a company you'd never heard of ten minutes ago. You should get straight answers before you paste a brief.

We don't store your prompts
Not for 30 days, not for 24 hours. We do not log, store or retain the content of your prompts or of model responses — so there is nothing on our side to leak, subpoena or delete. It is the first line of the privacy policy, not a setting you have to go and find.
What we keep instead
Metadata, because a bill has to come from somewhere: model, token counts, latency, HTTP status, cost, and a request id. That is the whole list, and it is what the usage table in your portal is built from.
What we will not claim
We are a router. Your request is proxied to the provider serving that model, and once it leaves us their handling applies — we have taken the training opt-out wherever a provider offers one, and we are not going to dress that up as a guarantee we are in a position to make.
Anyone telling you otherwise about someone else's infrastructure is guessing.

Start on us.

Free credits on the house when your account is created — enough to run your real workload against the models above and check the arithmetic yourself. No card, no contract, no sales call. Alpha seats open in small batches.

One email when your seat opens, with a login and a ready-to-run curl. No drip campaign, unsubscribe anytime.

Live system status Prompts and responses are never stored Prepaid credits — no card on file Sage AI LLC, a Wyoming company