AI Spend
What each spend figure means, who counted it, which regime it belongs to, and every honesty rule attached to it.
AI Spend
Nautir reports money carefully, because a spend figure that is wrong in the flattering direction is worse than no figure. Almost every honesty rule on this page exists because a plausible-looking number was once wrong by an order of magnitude.
Four questions sit behind every figure, and the product answers all four rather than collapsing them:
- Who counted it? DevZero's gateway, or the provider's own ledger.
- What kind of money is it? Real invoice dollars, or a notional price on prepaid traffic.
- What rate priced it? Your contracted rate, or a public list price.
- Against what was it measured? A control arm, or a before-and-after estimate.
Who counted it
| Provenance | Meaning |
|---|---|
| Provider-attested | The upstream's own ledger says this. Available only for DevZero-managed spend, and only once the usage sweep has read that period. |
| Gateway-reported | Measured by your own gateway as requests passed through. The only figure that can exist for spend on your own credentials, since DevZero never sees that provider's invoice. |
These are not two categories of spend. They are two ways of knowing what was spent, and the same arm can read gateway-reported in one period and provider-attested in the next, once the provider has been read for it.
The Spend screen splits your money into two arms -- Managed by DevZero and Your own provider keys -- and stamps each with its provenance.
Spend on your own credentials is for visibility rather than billing and may differ from what your provider charges you. DevZero never sees that invoice.
Why the daily bars do not foot to the total
Both daily series are gateway-measured, including the managed one. The provider's own figure is monthly-grained, so it cannot be split across days without inventing a distribution.
So the managed arm's daily bars show the shape of its spend while its section total shows the provider's figure. The page says so, rather than leaving you to notice the arithmetic.
Why a gateway total and a provider total legitimately differ
Tokens served from a cache never reached an upstream, so nothing was charged for them. The billing summary reports the cache-read token count and its share for exactly this reason -- it is the page's own explanation of the gap.
The two regimes read differently
| Regime | Credential | The money |
|---|---|---|
| Metered | Per-token API keys, yours or DevZero's | Real invoice dollars. |
| Subscription | Prepaid seats on OAuth-shaped credentials | Notional. Catalog-priced, with no invoice behind it. |
| Unclassified | No resolvable credential | Evidence of neither. Reported separately. |
Unclassified traffic is reported separately rather than folded into either side, because folding it left would pad the real-dollar figure and folding it right would inflate the notional one.
Every dollar figure on subscription traffic is notional. What the platform reports there instead is capacity: tokens not spent, and the extra headroom that buys inside a rate-limit window. Headroom percentages are estimates -- the vendor's exact token accounting for its caps is not published -- and they are labelled as estimates.
Seats are flat prepaid fees, so compression saves nothing there. The platform's role on seat traffic is to record utilization and attribution. Nothing on the seat surface is a cost-savings claim.
Contracted rates and list-price figures
| Basis | Meaning |
|---|---|
| Contracted | Your own negotiated per-model price, which you supplied. The authoritative input for any billable figure. |
| List price | The public price table. Directional only, labelled as such, and never billable. |
| Unpriced | No rate resolved. Reported as unpriced, never as zero. |
A list-price figure is systematically high for any discounted team -- and high in the direction that flatters the vendor. Wherever the product cannot resolve a contracted rate it says which models it fell back on rather than quietly mixing the two.
Two surfaces are explicitly billing-grade when your rates resolved end to end, and say so per response: the trial cost estimate and the experiment quote. A trial's running accrual is deliberately evidence-grade and list-priced -- read it as what the spend cap is enforced against, never as a charge.
And everywhere: unpriceable and free are different answers. A model whose rate could not be found is named, not priced at zero.
Where the money goes
The cost anatomy decomposes a provider's spend into components, and this is the single most useful thing to read before deciding what to optimize.
| Component | Notes |
|---|---|
| Uncached input | The only component prompt compression can address -- and roughly 1-3% of an agentic Anthropic bill. |
| Cache read | The dominant component on agentic clients. |
| Cache write | Anthropic-only. OpenAI and Gemini expose no cache-write counter at all, so this component is absent, not zero, there. |
| Output |
Read the decomposition's status before reading its components. A provider can be decomposed, or not decomposable -- and the reason is reported: no traffic, a model with no price, or rates that do not reconcile with what the traffic was billed at.
A component a provider does not have is absent, never a zero slice. The share denominator is always that provider's own decomposed total, never a cross-provider sum.
Evaluation spend is reported separately
Money spent on trials and experiments is reported beside your metered lines and is never summed into your total.
Every other figure is metered from request telemetry. This one is reported from the trial's own scores and never enters that table, because a trial row there would permanently raise a per-principal spend floor and later refuse a real user's request for an evaluation they had nothing to do with.
Folding a reported floor into a metered total would make the total neither.
It is a floor, and the reasons are echoed alongside it rather than left for you to discover. The window is half-open and both bounds are required: an absent bound would read as unbounded and render $0.00, which is indistinguishable from a team that ran no trials.
Advisory figures, and what they are not
Two read-only surfaces suggest where money might be recoverable. Both carry their own limits on the wire, because both are easy to over-read.
Model cost comparison reports what comparable work already cost on each model. It does not route, re-route, suggest a default, or change any served byte.
"Routine" work is not observable, so it is selected by a retrospective proxy -- short outputs, shallow conversations -- and the thresholds used are echoed back to you. A projection built on a small number of requests is a noisy rate, and the count is shown beside the projection for that reason. Any requests whose cost was estimated rather than provider-reported weaken every figure above them.
Upstream comparison compares what a marketplace charged against a published direct rate. Four things it cannot know are carried on the response:
- The two sides are not the same currency. The observed cost embeds the marketplace's margin; the counterfactual is a published list rate with no margin and no negotiated discount either. Their difference is an upper bound on recoverable margin, never a saving.
- There is no quantization dimension in the comparison.
- Going direct inherits work the marketplace does today: health measurement, failover, endpoint selection.
- Some cache rates are derived rather than published, and the response says when.
An unreported cost is not "charged zero". A reported $0 is an exact charge for a free model and counts as priced; a request the upstream reported nothing for is a different state, counted and never priced.
What to read next
- Savings -- measured versus estimated, and effective tokens.
- Efficiency -- waste still on the table, and who needs help.
- Attribution -- people, teams, departments, services and seats.