Skip to content
Journal / pricing

AI API billing transparency: what honest pricing looks like in 2026

What honest AI API billing looks like: per-call prices in every response, failed calls never billed, and a public fee you can check with a calculator.

I read pricing pages for a living. Running a tool router means our catalog tracks list prices across roughly 30 providers in 9 categories, and I have developed strong opinions about the difference between pricing that informs and pricing that obscures. This post is my attempt to write down the standard: what billing transparency should mean for AI APIs, and how to evaluate any vendor (including us) against it.

I have skin in this game, which I will get to. But the checklist stands on its own.

The four tests

1. Can you compute a price before you call?

The baseline test: given a request you are about to make, can you calculate what it will cost with a calculator and the public pricing page? For a lot of AI APIs the honest answer is no. Credits systems with shifting exchange rates, “contact us” tiers, per-feature multipliers buried in docs, units that do not map to anything you can measure in advance.

The per-unit pricing models in our catalog pass this test cleanly: search at dollars per 1k queries, parsing per 1k pages, transcription per audio minute, TTS per 1k characters. You can budget these. The failures tend to be multiplier schemes; LlamaParse’s premium mode consuming roughly 15x the base units is the canonical example in our catalog, not hidden exactly, but easy to miss until the invoice explains it to you. I covered the unit taxonomy in understanding per-unit pricing.

2. Do you learn the price when you pay it, or a month later?

Pre-call math is necessary but not sufficient, because in any system with routing, retries, or variable units, the actual cost of a call can differ from the naive estimate. The fix is simple and rare: put the exact price of the call in the response.

Every response from our router carries it: this call, this provider, this many dollars. Not a dashboard rollup, not an end-of-month surprise. The number arrives with the payload, so your own telemetry can attribute spend to sessions, tasks, and features in real time. Once you have worked this way, monthly-invoice-only billing feels like getting a restaurant check with no itemization. It also makes the whole system auditable: sum your responses, compare to your balance, and the two had better agree.

3. Do you pay for failures?

Here is a policy question that reveals a lot about a vendor: when a call fails, who eats it? A timeout, a 500, an empty result that should have been an error. If the answer is “you do,” the vendor’s incentives around reliability are quietly misaligned with yours.

Our rule is that failed calls are never billed. This matters double in a failover system: our router tries up to three providers per request, and the attempted-but-failed hops show up in the response’s routing.attempted chain for observability, but only the call that succeeded appears on the bill. You get the full story and pay for the ending. The transparency half of that design gets its own post in attempt chains and agent observability.

4. If there’s a middleman fee, is it a number?

Any aggregator, router, or platform between you and providers takes a cut. The transparency question is not whether the cut exists; it is whether it is published as a checkable number or diffused into opaque bundle pricing where you cannot tell what the underlying service costs.

This is where my skin in the game lives, so let me state ours plainly: route.tools charges provider list price plus 20%, on everything, published at /docs/pricing. Serper lists at $1.00 per 1k queries; we charge $1.20. You can verify any line in our catalog against the provider’s own pricing page, and I would encourage it, because the ability to check is the entire point. If the 20% offends you at your volume, the BYOK option drops to a 5% platform fee on your own provider keys, also a public number. I am not claiming these are the smallest possible fees. I am claiming you can compute them, and that this property is rarer than it should be.

Why this matters more for agents

Human-driven API usage has a natural governor: a person notices weirdness. Agents do not. An agent will happily retry a failing call in a loop, or hammer a mispriced endpoint all night, and a billing system that only reports monthly turns those incidents into four-figure surprises. Per-call prices in responses, combined with prepaid credits and spend caps, form the layered defense: you see spend as it happens, and there is a hard ceiling when you miss it anyway. I wrote about the ceiling half in spend caps for AI agents.

The checklist, portable

Evaluating any AI API or platform, ask:

  • Can I compute a request’s cost in advance from public numbers?
  • Does the response tell me what the call actually cost?
  • Are failures billed?
  • Is every middleman fee a published number?
  • Can I reconcile per-call prices against my balance, myself?

Vendors that pass all five are betting their pricing survives scrutiny. Vendors that fail several are betting it will not get any. That difference tells you most of what you need to know.

If you want to see per-call pricing in the wild, signup comes with $2 in free credits, enough for a few thousand calls in the cheaper categories: dashboard.