You thought you bought a subscription. You bought a meter.
ChatGPT, Claude and Copilot all now bill on usage: your seat comes with an included amount, and above that a meter runs. This page explains how that billing works, why it catches people out and what to do about it.
What are usage credits on AI subscriptions?
Usage credits are a balance you buy on top of your AI subscription. Your seat includes a fixed amount of usage; once that's gone, work either stops or continues from a credit pool billed separately. ChatGPT, Claude and Microsoft Copilot all now work this way. A fixed monthly figure has quietly become a variable cost.
- A seat includes usage; above it a meter runs
- Every provider has its own, non-comparable unit
- Limits run out mid-week, not on the invoice date
- One admin setting turns a fixed cost variable
From subscription to meter — and almost nobody saw it coming
AI subscriptions started out as something familiar: an amount per user per month, like any other software package. You knew what it cost, you knew what you got. That era is over, and the transition happened so gradually that most organisations missed it.
What replaced it is a hybrid: you still pay per user, but that seat includes a set amount of usage. Above it a meter starts running — credits at ChatGPT and Copilot, usage credits at Claude — billed separately. The subscription is no longer a ceiling but a floor.
As long as people use AI as a chat window, you notice little of this. The included amount is generous enough for asking questions and drafting text. It tips the moment someone runs deep research, hands an agent a task or has a stack of documents processed — then a multiple goes out per action, and that is exactly the usage you want as an organisation.
Hence the surprise: the bill climbs because of precisely the people who are best with AI. That's no reason to slow them down. It's a reason to make sure they know what they're turning.
How the billing works
Three major providers, the same principle, three different implementations. The details change regularly — the mechanics beneath them don't.
ChatGPT
An included limit, then a shared credit pool
- Every seat comes with an included amount of usage.
- Beyond that you draw from a workspace credit pool the admin buys upfront.
- Heavy features — deep research, thinking models, image generation, voice, coding agents — all draw from that same pool.
- Purchased credits carry a twelve-month shelf life.
Claude
Limits that reset at a fixed point each week
- Seats come with included usage and a weekly limit.
- That limit resets at a fixed time assigned to your account — not on the first of the month.
- At the limit you can wait for the reset, move to a heavier plan, or the admin switches usage credits on.
- With those on, work continues and the extra usage is billed separately.
Microsoft Copilot
Credits, prepaid or after the fact
- Agents and Copilot Studio bill in credits — until September 2025 these were called 'messages'.
- You either buy prepaid capacity packs or settle afterwards through Azure.
- Different actions consume different numbers of credits.
- Anyone already holding a Microsoft 365 Copilot licence is partly covered by it.
Three providers, three units, no shared meter reading. Anyone using two of the three has two different ways of counting — and not a single screen that adds them up.
Why this catches people out
Not because it's hidden anywhere — it's right there in the terms. But because it behaves differently from all the software you bought before it.
01
The unit is invented — and not comparable
A credit is not a token is not a message is not a premium request. Every provider invents its own unit, sets its own exchange rate and adjusts it: Microsoft renamed 'messages' to 'credits'. So you can't put two offers side by side on the unit — only on what you can actually get done with it.
02
Not every action costs the same
Asking a short question and launching a deep research run look almost identical in the same interface. In consumption they differ by an order of magnitude. Nothing on the screen tells the user which of the two they just did.
03
The limit hits mid-task
Not on the invoice date, but on Thursday afternoon for the person who has to finish the report. That's why this lands as a surprise: it isn't an accounting problem that arrives at finance, it's a work stoppage for someone with a deadline.
04
Balances can expire
Prepaid credits have a shelf life. Buying generously to be rid of the hassle is therefore not a free choice — too much is wasted money just as surely as too little is a work stoppage.
05
One checkbox turns fixed into variable
An admin switches on 'keep working past the limit' to help a colleague who is stuck. From that moment the subscription is no longer a ceiling but a floor, and for the rest of the year the question isn't 'what does it cost' but 'what did it come to'.
Why consumption climbs so fast
Beneath every credit sits the same unit: tokens, the pieces of text a model counts in. Four things explain why so many of them go through.
Context travels along at every step
A model remembers nothing between two calls. All instructions, retrieved documents and earlier answers go along again each time. Handle a task in twenty steps and you pay for that context twenty times.
Agents take dozens of steps per task
Plan, look up, consult a system, verify, adjust. Analysts measure agentic tasks consuming a multiple of a single chat question — and how large that multiple is depends on what the model decides along the way.
What comes out weighs more than what goes in
Output counts several times heavier than input at virtually every provider. Models that think out loud first also produce intermediate steps you never see but do consume.
Failed attempts count too
A call that gets retried, an agent circling, a user rephrasing three times. None of it is logged as an error — it just disappears into the total.
More on how agents work and why they consume so much is in the knowledge base under AI agent and RAG.
What to do about it: literacy, awareness, standardisation
The answer to consumption costs is rarely a dashboard. It sits in what people know, what they can see and how firmly your way of working is fixed.
01
Literacy
People can't be careful with something whose unit they don't understand.
- Knowing what a credit is and what draws from the pool in your situation.
- Recognising when you're setting an agent to work rather than asking a question.
- Knowing there's a limit, when it resets and what to do when you hit it.
- Understanding that dragging a long conversation along costs more than starting fresh.
This isn't a finance topic, it's a user skill. Half a day of explanation for the people who open the tool daily does more than any dashboard.
02
Awareness
Consumption that only becomes visible on the invoice was seen a month too late.
- Every application has an owner who knows what it consumes.
- Consumption is visible in the moment, not only afterwards as a total.
- An alert when the pattern deviates, not just when the ceiling is reached.
- A short monthly look back: what did we run, what did it deliver.
Whoever can see, can steer. In practice consumption drops the moment people know somebody is watching — before a single agreement has been made.
03
Standardisation
Consumption only becomes predictable once the work beneath it is predictable.
- One tool per job, instead of three assistants doing the same thing.
- An agreed default model, with a reason required to deviate.
- Fixed working methods and shared prompts for work that keeps coming back.
- New application? It passes by someone before it goes live.
Without a standard, consumption depends on whoever happens to pick up the job. With one, you know roughly what a process costs — and a deviation stands out immediately.
SME or enterprise: the gain sits in a different place
The billing works the same, the biggest saving does not. Start where your organisation is leaving the most on the table.
SME
Clean up what you already have
- Licences you pay for people who barely open the tool.
- Three assistants side by side doing largely the same work.
- Automations consuming in the background without an owner.
- One person keeping an eye on consumption, instead of nobody.
Usually a matter of sessions, not months. The gain is in the overview: knowing what you buy, who uses it and what it returns.
Enterprise
Steer on what is spread out
- Attributing consumption to teams and applications instead of one collective line item.
- Ceilings and alerting in the software, not only in policy.
- A cost check before a new application goes into production.
- Agreements on which model is the default and when you escalate.
Here the individual intervention isn't the problem — the problem is that nobody sees the whole. Control appears the moment every team sees its own consumption and is held to it.
How we approach it
Not a clean-up operation, but four steps that make consumption visible and keep it that way.
-
Take stock of what you buy
first stepWhich subscriptions are running, which credits were bought, who uses what and what is set to auto-recharge. At virtually every organisation this step alone turns up surprises.
-
Make agreements that hold up
short lead timeWhich tool for which job, which model is the default, when you escalate and what passes by someone before it goes live. Short and concrete, not a policy document.
-
Bring people along in their own work
per teamExplain what consumes and how to get the same result more lightly, using the tasks they actually do. Not general AI training, but their own work.
-
Keep it visible
ongoingAn owner per application, ceilings where possible and a short monthly look back. That way control becomes a habit instead of a one-off clean-up.
Do you know what you're currently buying?
The AI Readiness Scan maps which tools and applications are running, who uses them and where you lack control. A concrete starting point for the conversation about cost.
Take the AI Readiness ScanFrequently asked questions about usage credits
What exactly are usage credits?
Usage credits are a balance you buy on top of your AI subscription. Your seat includes a fixed amount of usage; once that's gone, work either stops or continues from a credit pool that is billed separately. ChatGPT, Claude and Microsoft Copilot all now work on that principle, each with its own unit and its own rules.
Why do I hit a limit when I'm paying for a subscription?
Because the subscription doesn't cover unlimited use, but an included amount. That distinction goes unnoticed for a long time: with normal chat use you rarely reach it. The moment someone runs deep research, deploys agents or has large documents processed, that amount suddenly is the binding constraint — and you find out mid-task, not on the invoice.
Can I compare providers on their credits?
Not directly. A credit at one is not a credit at another, and the provider decides how much each action draws and can change it. So compare not on the unit but on what you can get done in your own work: how much of a typical process fits inside what you buy, and how often you ran into the limit over recent months.
What happens when the credits run out?
That depends on what's been configured. With nothing switched on, usage stops until the limit resets or the admin buys more balance. With continue-past-the-limit on, work carries on and the extra usage is billed separately. The first gives you a ceiling with work stoppages, the second no stoppages but also no ceiling.
So should we switch that continue option on or not?
Usually yes, but not without agreements alongside it. People stuck mid-task cost more than the consumption you save. Switch it on, but with an owner watching, a ceiling where that's possible and the agreement that deviating consumption gets discussed. Switch on and forget is the variant the surprise comes from.
Why is our AI bill rising while per-token rates are falling?
Because consumption grows faster than the price drops. Applications get better by sending more context: longer instructions, retrieved documents, conversation history, intermediate steps. Every improvement costs consumption. You only feel the cheaper unit once you also steer on the number of units you use.
Is this relevant for a smaller company too?
Yes, but the attention goes elsewhere. In smaller organisations the money usually sits in licences nobody uses, in three assistants side by side doing largely the same work, and in automations consuming in the background without anyone watching. That is often cleared up in a few sessions, whereas larger organisations are more likely to face consumption spread across many teams.
How do I stop it getting out of hand again in six months?
By not treating it as a one-off clean-up. An owner per application, a short monthly look back and the agreement that a new application passes by someone first. Combine that with explanation for the people using the tools daily; they make a hundred small decisions a week that together determine your bill.
Is your AI bill climbing faster than you expected?
Book a no-obligation conversation. We look at what you buy, where the consumption comes from and which agreements will have the fastest effect in your situation.


