Skip to content

The AI cost squeeze is already here, just not where you're looking

Real numbers 10 min read

The popular framing goes something like this: AI services are heavily subsidised, the labs are losing huge sums of money, and at some point that party has to end. ChatGPT Plus jumps from £20 to £200, the API rate card doubles, and everyone wakes up to a new, more expensive reality.

That story is half right. The squeeze is real. But it is already happening, and it is not being delivered through the front door.

Sam Altman publicly admitted in early 2025 that OpenAI was losing money even on its $200/month Pro tier because power users were consuming more than the price assumed. OpenAI's own projections do not have the company turning a profit until 2030. Anthropic is in better shape, but is still burning a few billion a year, against a revenue run-rate that hit roughly $30 billion in spring 2026. The economics underneath the polished consumer experience are not stable.

What is interesting is what the labs and their distributors are doing about it. They are not raising the sticker price. They are tightening the screws in places most users don't think to look.

What the squeeze actually looks like

A few patterns to watch for, because they affect your bill more than the headline price does.

Tier creep. In about 18 months we have gone from a £20 Plus subscription as the top consumer tier to £100, £200, and now £300 plans. The most capable models have quietly migrated upstairs. The £20 tier today is roughly what £200 bought you a year ago. That is not a price hike, it is the goalposts moving.

Metering changes. Token-based billing is starting to expand into places that used to be flat-rate. GitHub Copilot is moving all subscribers to usage-based billing from June 2026, with monthly credits that run out and overage charges thereafter. That is a meaningful change in unit economics for any team that was comfortable with predictable seat pricing. Tokenizers themselves can be tweaked, so the same English sentence can suddenly cost a few percent more tokens to send and receive than it did last month.

Free tiers being deliberately diluted. ChatGPT Free and Go now serve ads in the US and selected markets, having launched in February 2026. The free tier has not disappeared, but it is being reshaped to push you toward paid tiers, and what you put into a chat with an ad layer over the top is worth thinking about.

Agentic workflows quietly compounding token use. This is probably the most important shift, and the one I see least understood. A single agent task, where the model is reading tools, writing tools, taking actions, then re-reading everything before the next step, can consume 5 to 30 times the tokens of an equivalent chat. Some studies put it much higher than that. The per-token price keeps falling, but the per-task price is not, because the task got bigger.

Net effect: per-token prices keep dropping, while monthly bills keep climbing. a16z's research shows average enterprise AI spend rising from around $4.5 million to $7 million in two years, with leaders expecting another ~75% growth in the year ahead.

So the question is not really "will prices rise". The question is "how do I run my AI use in a way that does not get more expensive every time the underlying market moves?"

What to do

A handful of habits genuinely help. None of these are revolutionary, they are just rarely treated with the same discipline you would apply to hosting or SaaS.

  1. Avoid lock-in. This is the single most important one. If your prompts, workflows, and integrations only work on one provider, you have removed your own ability to switch when something changes. Keep your prompt library in plain text or markdown, not in custom GPTs or provider-specific projects. If you are building automations, route AI calls through a tool like Zapier or Make where you can swap the model behind the scenes, rather than hard-coding to a single API. Test the same prompt across two or three providers occasionally so you know what works where.

  2. Right-size each task. Most of what people are putting through the most capable model on offer does not need it. Take Anthropic's Claude range as an example: Sonnet is the everyday workhorse, fine for drafting, summarising, idea generation, and most coding work. Opus is meaningfully better on harder reasoning and complex synthesis, and it can feel the difference on long-context work too, but it costs more per token and is slower. The sensible default is to start in Sonnet and reach for Opus only when you can see Sonnet struggling. The same shape holds elsewhere: GPT-5.4 mini for everyday work and GPT-5.5 for harder problems on the OpenAI side, Gemini 3 Flash and Gemini 3.1 Pro on Google's. Building the habit of picking the right model for the job is more useful than picking a favourite tool.

  3. Track the changes. Tokenizer changes, deprecations, tier reshuffles, model retirements, billing model changes. Subscribe to the changelog or release blogs of the providers you actually use. Set a calendar reminder once a month to have a look at what has shifted. Most of the people I have seen surprised by their AI bill were not paying attention to the small print updates.

  4. Audit the spend. Once you are running more than one AI tool, it is easy to end up with three or four subscriptions on the books, two of which someone signed up for during a pilot last year and then quietly forgot. Tag each subscription against the work it actually supports. If you cannot name the workflow, cancel it.

  5. Lock in annual billing on the one tool you are sure you will use for the next year. It is a small hedge, usually saving 10 to 20%, and it protects you if the monthly price moves before your renewal.

What not to do

The mirror image is just as useful.

  1. Don't put AI into a workflow just because you can. AI is genuinely useful, but it is not free, it is not deterministic, and it is not always the cheapest tool for the job. A regex, a SQL query, or a five-line script will often do what people are asking a language model to do, more reliably and at zero marginal cost.

  2. Be deliberate about what you put into ad-supported tiers. Now that ads sit on top of the consumer free and Go tiers, what you type is shaping what the ad layer thinks it knows about you. That makes free tiers the wrong place for client data, anything covered by an NDA, and your own commercially sensitive information too. If you would not paste it into a Google Doc you did not own, don't paste it into a free chatbot.

  3. Don't bet your unit economics on today's prices. A few years back I built a small service with a friend that wrote short messages onto a public blockchain. The maths worked, until it didn't. The cost of writing each message rose past what we could plausibly charge for it, and putting our prices up enough to fix it would have killed the demand we did have. The product quietly stopped being viable. AI features carry the same shape of risk. If you are building something that uses AI in production, design it as if costs could be several times higher than today, with hard per-customer ceilings and the option to swap the model behind any given feature.

  4. Don't put your only knowledge base inside one provider's product. Custom GPTs, Claude Projects, and Gemini Gems are genuinely useful, and they get more useful the longer you use them, because they pick up on how you work and what you care about. The downside is the same coin: that learning sits in the provider's account, and it does not come with you if you switch or they change the product. The pragmatic answer is to use them for what they are good at, but keep a parallel, portable version of your context (plain markdown in Notion, Obsidian, or a folder you own) so you are never wholly dependent on one provider's memory of you. I am working on something tidier here, more on that another time.

When AI isn't the answer

Probably the most useful frame I have found, and the one that quietly does more for the bill than any of the others, is this: AI is excellent at finding the rule. It is often the wrong tool to run the rule. It is also, conveniently, pretty good at building the thing that runs the rule.

For example, you have a recurring finance reconciliation. Maybe it is matching invoices to bank transactions, or pulling line items out of supplier statements and checking them against your accounting system. The first time you do it, AI is genuinely useful. You can paste in a sample, talk through the edge cases, get help spotting patterns, and end up with a clear specification of what you actually want to happen.

Once you have the rule, the next conversation is "now please write me the Zapier flow, the Google Apps Script, or the small Python script that does this every month." Modern models are surprisingly capable here, and once it is built, the script runs on a schedule for free, with no token cost and no provider risk attached.

I have seen people end up paying ongoing API charges to summarise the same kind of document week after week, when the underlying logic could have been pinned down once and handed to a deterministic tool. The reverse pattern is usually cheaper, more reliable, and unaffected by anyone's pricing changes.

This is not a rule, it is a question. When I find myself reaching for an AI tool repeatedly for the same task, I now ask: am I still discovering, or am I just executing? If it is the latter, there is probably a small tool waiting to be built that takes AI out of the loop.

A quieter takeaway

The people I have watched come through the last 18 months in good shape were not necessarily the ones who picked the right model. Most of them changed providers at least once. They were the ones who built the habit of asking, before reaching for an AI tool, whether they actually needed one for this particular job, and whether what they were about to build would still work if the provider behind it changed its mind.

That is the discipline worth practising. The price of a million tokens will keep falling, your bill will keep rising, and the right answer will keep being some combination of fewer tokens, cheaper tokens, and fewer tasks that need tokens at all.


Sources: Sam Altman on ChatGPT Pro losing money (Fortune, Jan 2025); Anthropic revenue run-rate (Bloomberg, Mar 2026); GitHub Copilot moving to usage-based billing (GitHub Blog, 2026); ChatGPT advertising rollout (OpenAI, Feb 2026); Agent token consumption (Stanford Digital Economy Lab); a16z AI Application Spending Report; LLM inference price trends (Epoch AI).