What happened
Over the past few weeks all three major model providers refreshed their pricing and context windows, and the headline is the same everywhere: cheaper, with more context. Sonnet 5 arrives at an introductory two dollars per million input tokens and ten for output, cheaper Gemini variants drop below two dollars, and context windows have grown to a million tokens and beyond. If the only number that mattered were the top line of the price sheet, the story would be closed: you get more for less.
The trouble is that the bill stopped tracking that top line. Underneath, pricing has fractured along several axes at once — and each one can move your cost independently of the advertised rate.
The first is the long-context tier. With one provider, crossing roughly two hundred thousand tokens in a request moves the whole request onto a pricier schedule — input goes from two dollars to four, output from twelve to eighteen. With a second, long context means doubled input and markedly higher output than short prompts. With a third, the full million-token window is billed at the same rate as a short prompt — a nine-hundred-thousand-token request costs the same per token as a nine-thousand-token one. The same "big memory," then, has three entirely different cost curves depending on whose model you wired in.
The second axis is region and data residency. All three charge a premium of roughly ten percent to route inference through a guaranteed geography, or to keep it inside a given country. The default "global" rate — the one in the headline — is not the price paid by anyone who, for GDPR or AI Act reasons, has to keep processing in Europe. The compliant configuration is the more expensive one.
The third axis is the least visible: the tokenizer. One provider's newest models count the same text at roughly a third more tokens than their predecessors. The per-token rate hasn't changed — but the same document and the same conversation are now a third larger bill if you switched to a newer model "at the same price."
The fourth is plain effective dates. Sonnet 5's introductory price ends on 31 August 2026 — from 1 September the rate rises from two dollars to three for input and from ten to fifteen for output. That's not a surprise increase; it's an increase already on the calendar, worth putting into the forecast now.
Our thesis
The top line of the price sheet is the least useful number you have today. "AI got cheaper" is true on the label and false on the invoice — because the bill is driven not by the advertised rate but by which of these axes moves next, and whether you're hard-wired to the provider that moved it. We wrote about the same concentration risk when Copilot moved to usage-based billing — only now it sits a floor lower, on the model layer itself rather than the tool.
We expect the number of axes to grow, not shrink — that's our read of the market, not an announcement from any provider. The more dimensions a price sheet has, the less sense it makes to optimize for a single rate, and the more sense it makes to hold one architectural assumption: no single provider is wired into the product permanently.
Why it matters
Private Equity
For a fund the dependency is simple: if a portfolio company built its margin thesis on one provider's rate today, part of that thesis is somebody else's pricing decision. The point isn't to guess who will raise prices — it's whether the company can change provider at all without rewriting the product. That's a due-diligence question alongside the compliance ones: show the plan B for the model, not just which model you use today.
Enterprise
For a large organization these axes converge in one place. If you have to keep inference in Europe, you pay the regional tariff, not the global one. If your requests — typically in document-heavy solutions — cross the long-context threshold, with some providers you hit the second, pricier tier and with others you don't. The conclusion isn't "pick the cheapest," it's "don't wire yourself to one provider so tightly that its change to a threshold, a tokenizer or a region becomes your product rewrite." In practice that means a layer that routes each task to a good-enough rather than the most expensive model — and lets you swap a provider by configuration, not by project.
SMB / mid-market
A smaller team usually picks one model and stays with it, so it gets hit by the least visible axis: you switch to a newer variant "at the same price," and the bill rises because the new tokenizer counts your text more expensively. The good news is you don't need a FinOps department for this — it's enough to check once what a typical request actually costs after a model change, before you take the saving as a given.
One step you can take this week
Take one AI process you already run in production and describe it in four lines: which provider and model, whether requests regularly cross the long-context threshold, whether you have to keep processing in a specific region, and what it would actually take to move that process to a different provider. If the answer to the last one is "rewrite the product," then that cost — not the per-token rate — is your real cost of AI inference. The intermediary layer that switches models and locations without touching the product is what we describe under products.
Describe your case
If you pay for a single model, hard-wired, and don't know what it would cost to change providers or to keep processing in Europe, bring one process and the name of the model it runs on. We start from something concrete. Describe your case: mailto:[email protected]?subject=Rozmowa%20z%20Aurora%20AI.