AI & Automation
The key question about AI token usage: expensive, or very expensive?
Why is the cost of AI token usage getting more and more attention? Yes, it comes down to money. Because we love using AI, but we may be slightly fonder of our budgets...
András Leskó
CTO

At Gloster—and I imagine at your company as well—AI is becoming increasingly integrated into our daily development and business work, typically in so-called “agentic” workflows based on Claude Code and Microsoft Copilot, either under human supervision or, in some cases, running autonomously for extended periods. As usage grows, it’s becoming harder and harder to avoid the question: how much does all this cost, and how can we keep costs under control? And above all, how do we use it so that it’s worth it? The problem is already here, and it’s becoming increasingly painful. During an internal kick-off discussion, we explored this from several angles: how can we manage AI token usage so that we don’t disrupt daily work, yet also ensure the cost of using AI doesn’t spiral out of control? It turned into a lively debate, but in the end, we reached a common position.

Where are prices headed?

To get a realistic view of the issue, it’s worth looking at market trends first.

The trend is moving toward a usage-based model. In recent months, a pattern has emerged in which major providers are increasingly shifting to usage-based pricing, and in the case of Microsoft Copilot, this is already clearly evident today. If this trend continues and spreads to all providers and all their models, sooner or later everyone will have to reckon with the pay-as-you-go model. So it’s worth not only panicking after the first bill arrives, but also preparing for it in advance.

It matters what we use AI models for. One-time question-and-answer tasks are inexpensive. Agent-based processes, on the other hand—where the model works autonomously in multiple steps—consume orders of magnitude more tokens: every iteration, every file read, and every redesign places a burden on the framework. As we shift an increasingly larger portion of our daily work to such processes, consumption does not increase linearly but rather in leaps and bounds. That is precisely why it is not enough to simply say, “We use AI.” What matters most is what we use it for, how much we use it, and how efficiently we use it.

Points of View

Imposing premature restrictions on developers is counterproductive. The technical team’s position is that we must first learn to use the tools effectively, and only then should we optimize costs. It’s a long-standing truism in software development that premature optimization is the death of a project, and this holds true here as well. If we impose tight constraints too early, we stifle the very learning phase when the team is just beginning to discover where the tool delivers the most value. It’s worth embarking on cost optimization only once the organization has achieved the necessary AI maturity.

The pay-as-you-go pricing model is flexible, but less predictable for customers. In this model, limits are typically set at the workspace level; however, this makes it harder to plan usage, and even with nominal limits in place, there may still be surprises on the bill. This alone underscores the need to plan, monitor, and control actual usage; otherwise, you won’t know the final cost until the bill arrives.

A pre-purchased budget is predictable, but if it runs out early, it can easily disrupt daily operations. The advantage of license-based models, such as a team license, is that the budget is fixed in advance, and for many teams, this is sufficient most of the time. But if you ask your sales manager, they’ll surely tell you that the daily token quota will run out at the worst possible moment—for example, during the final hour of an urgent bid submission. And by the time the quota is replenished, the deadline will have passed, you’ll be out of the bid, or the issue will have long since become irrelevant to the client. And referring back to market trends, it’s also unclear how long these pricing structures will remain in place at all.

Perhaps not everyone needs every model. The technical team has noted that not every department within an organization necessarily needs the most powerful model. For finance, for example, a more economical model may suffice, and access is naturally regulated from the outset without the need to introduce separate restrictions. We are not advocates of wrapper solutions, but differentiation based on the model layer is a realistic approach: choosing a model tailored to the task simultaneously reduces costs and keeps usage within reasonable limits.

The common denominator: we need to measure

By the end of the conversation, we agreed on one thing: No matter which direction we take—whether it’s pay-as-you-go, a license-based framework, or differentiation by model tier—without measurement, we’re just groping in the dark.

Measurement shows who has actually adapted, in which areas and for what tasks resources are being used, where there is waste, and where the return on investment is realistic. Without this, any restrictions—and any rejection of them—are based solely on gut feelings; it’s just a lottery, roulette, or a trap. What we don’t measure, we can’t control.

At Gloster, we no longer view this as a purely theoretical question: we’ve started measuring and optimizing our own use of AI. We monitor, record, and analyze which teams use what and how much, where a more powerful model is worth it, and where a more economical one suffices—and from this, we’re gradually developing best practices for AI usage. Our goal is, on the one hand, to save a ton of money (because, as we’ve said, we love that…), and on the other hand, to point the way forward and share along the way what works for us and what we recommend others try and implement. Because this issue affects you, too—and everyone else who relies on AI to boost their company’s efficiency. So, pretty much everyone.

So what happens now? Over the next few weeks, we’ll break down the issue at hand and the possible solutions in separate articles, organized by role. The main topics—and what’s already taking shape—are:

  • Technical and Delivery Perspective: How can token usage be measured, and what conclusions can be drawn from this?
  • Sales and Market Reality: What We See in the Market. This is the result of the shift toward usage-based models and the downward trend in token prices.
  • Business and executive perspective: Where does all this pay off financially, and what are the real pain points at the CEO level?
  • Conscious Use of AI: Which Tools and Models to Use for What. We’ll also share insights from our own Hermes agent here.
  • Training and HR: How to Help Developers and Non-Developers Use AI Correctly and Effectively.

Let's continue!

At Gloster, and I imagine at your company too, AI is being built ever more deeply into everyday development and business work, typically in Claude Code and Microsoft Copilot based workflows: sometimes under human supervision, sometimes already running on their own for long stretches, in what are now called "agentic" setups.

As usage grows, one question becomes harder and harder to avoid: how much does all this cost, and how can those costs be kept under control? And above all, how do we use it so that it actually pays off?

The problem is already here, and it is getting more painful. We looked at it from several angles in an internal kick-off discussion: how do we manage AI token usage in a way that does not hold up the daily work, but also does not let the cost of using AI run away from us? It sparked a good debate, but in the end we did arrive at a shared position.

Where the pricing is heading

Before we argue about control, it is worth reading the market.

The direction of travel is consumption-based. Over the past few months the picture has sharpened: the large suppliers are moving towards usage-based pricing, and with Microsoft Copilot you can already see it in the open. If that continues, spreading across every supplier and every model, pay-as-you-go becomes the default everyone plans around. So the smart move is to prepare for it now, rather than panic when the first bill lands.

What you use AI models for makes all the difference. A single question-and-answer exchange is cheap. Agent-based workflows — where the model works across many steps on its own — burn tokens on another order entirely: every iteration, every file it reads, every replan draws down the budget. As more of the daily work shifts into those workflows, consumption does not climb in a straight line. It jumps.

Knowing that "we use AI" tells you almost nothing. What counts is what you use it for, how much, and how well.

The competing views

Fencing developers in too early is counterproductive. Software engineering has an old truth that premature optimisation kills the project, and it holds here too. Draw the budget tight too soon and you smother the exact learning phase where the team is still finding where the tool earns its keep. Cost optimisation is worth starting once the organisation has the AI maturity to support it.

Pay-as-you-go is flexible, but harder for the buyer to predict. In this model you can usually set limits at workspace level, and in return consumption is tougher to forecast. Even with nominal limits in place, the billing can still surprise you. That alone makes the case for planning, monitoring and controlling actual usage.

A pre-purchased allowance is predictable, but it can stall the day's work if it runs out early. Licence-based arrangements fix the budget up front, and for many teams, most of the time, that is enough. But the daily token allowance can run dry at the worst possible moment — in the final hour before an urgent bid goes out. And it is an open question how long these arrangements survive the market shift at all.

Perhaps not everyone needs every model. Not every function in an organisation needs the strongest model. Finance, for instance, may be well served by a leaner one, and that choice regulates access naturally, without imposing a separate cap. Matching the model to the task cuts the cost and steers usage back into sane territory at the same time.

The common ground: it needs to be measured

By the end of the conversation we agreed on one point. Whichever way we go — pay-as-you-go, a licence-based allowance, or differentiation by model tier — without measurement we are feeling our way in the dark.

Measurement shows who has genuinely adopted, which areas and which tasks the consumption goes on, where there is waste and where the return is real. Without it, every limit — and every decision to lift one — rests on nothing but a hunch.

What you do not measure, you cannot keep in hand.

So at Gloster we no longer treat this as theory. We have started measuring and optimising our own AI usage. We monitor, log and analyse which team spends what and on what, where the stronger model pays for itself and where the leaner one is enough, and from that we are building our best-practice AI playbook step by step.

Our aim is partly to save a serious pile of money, and partly to point the way, and to share as we go what works for us and what we would suggest others try out and adopt. Because this problem affects you too, and everyone else who is building their company's efficiency on AI. So, pretty much everyone.

Over the upcoming weeks we will unpack the problem and the possible solutions role by role, in separate articles.

To be continued!

What is coming next — the series ahead

  • Technical and delivery view — how token usage can be made measurable, and what conclusions follow from it.
  • Sales and market reality — what we see in the market: the shift towards usage-based models, and the net effect of the downward trend in token prices.
  • Business and leadership view — where all this pays back financially, and the real pain points at CEO level.
  • Deliberate AI usage — which tool and which model to bring to which job, including what we have learned from our own Hermes agent.
  • Onboarding and HR — how to help both engineering and non-engineering colleagues use AI correctly and efficiently.

Bei Gloster, und vermutlich bei Ihnen auch, wächst KI immer tiefer in den Entwicklungsalltag und in die Arbeit der Fachbereiche hinein. Typischerweise in Workflows auf Basis von Claude Code und Microsoft Copilot, unter menschlicher Aufsicht – manchmal aber auch schon in agentischen Workflows, die über längere Zeit selbstständig laufen.

Je stärker die Nutzung steigt, desto schwerer lässt sich die Frage umgehen: Was kostet das alles, und wie behält man die Kosten im Griff?

Das Problem ist längst da, und es tut zunehmend weh. In einem internen Kick-off haben wir es aus mehreren Blickwinkeln beleuchtet: Wie gehen wir mit dem Tokenverbrauch um, ohne die tägliche Arbeit auszubremsen und ohne dass die Kosten davonlaufen? Daraus wurde eine ordentliche Debatte – am Ende aber auch ein gemeinsamer Standpunkt.

Wohin steuert die Preisgestaltung?

Der Trend geht in Richtung verbrauchsabhängiger Modelle. In den letzten Monaten hat sich abgezeichnet, dass die großen Anbieter zunehmend auf verbrauchsbasierte Preise umstellen; bei Microsoft Copilot ist das heute schon konkret sichtbar. Setzt sich diese Richtung fort und erfasst sie jeden Anbieter und jedes Modell, muss früher oder später jeder mit Pay-as-you-go rechnen.

Es lohnt sich also, nicht erst nach der ersten Rechnung in Panik zu geraten, sondern sich vorab darauf einzustellen.

Es macht einen Unterschied, wofür wir die Modelle einsetzen. Einzelne Frage-Antwort-Interaktionen sind günstig. Agentische Abläufe dagegen, bei denen das Modell in mehreren Schritten selbstständig arbeitet, verbrauchen um Größenordnungen mehr Token: jede Iteration, jede eingelesene Datei, jede Neuplanung belastet das Budget. Je größer der Anteil solcher Abläufe an der täglichen Arbeit wird, desto stärker steigt der Verbrauch – nicht linear, sondern sprunghaft.

Deshalb genügt es nicht zu wissen, dass wir KI nutzen. Entscheidend ist, wofür, wie viel und mit welcher Effizienz.

Die Blickwinkel aus dem internen Kick-off

Entwickler zu früh einzuschränken ist kontraproduktiv.Zuerst muss man lernen, die Werkzeuge gut zu nutzen, erst danach optimiert man die Kosten. Vorzeitige Optimierung erstickt genau die Lernphase, in der das Team herausfindet, wo das Werkzeug den größten Nutzen bringt.

Pay-as-you-go ist flexibel, für Kunden aber weniger planbar.Limits lassen sich auf Workspace-Ebene setzen, doch der Verbrauch ist schwerer zu planen, und selbst bei nominellen Grenzen können in der Abrechnung Überraschungen auftreten.

Ein vorab gekauftes Kontingent ist planbar – geht es zu früh zur Neige, blockiert es aber die tägliche Arbeit.Das tägliche Tokenkontingent geht im schlechtesten Moment zur Neige, etwa in der letzten Stunde vor einer dringenden Angebotsabgabe.

Vielleicht braucht nicht jeder jedes Modell.Innerhalb einer Organisation benötigt nicht jeder Bereich das stärkste Modell. Die Wahl des passenden Modells je Aufgabe senkt die Kosten und lenkt die Nutzung zugleich in vernünftige Bahnen.

Der gemeinsame Nenner: Man muss messen

Am Ende des Gesprächs waren wir uns in einem Punkt einig. In welche Richtung wir auch gehen – Pay-as-you-go, lizenzbasiertes Kontingent oder Differenzierung auf Modellebene –, ohne Messung tappen wir im Dunkeln.

Die Messung zeigt, wer die Werkzeuge tatsächlich nutzt, in welchem Bereich und für welche Aufgaben der Verbrauch anfällt, wo Verschwendung entsteht und wo sich der Einsatz realistisch rechnet.

Was wir nicht messen, können wir auch nicht im Griff behalten.

Bei Gloster betrachten wir das deshalb nicht mehr als theoretische Frage: Wir haben begonnen, unsere eigene KI-Nutzung zu messen und zu optimieren. Wir erfassen, speichern und analysieren, welches Team wofür und wie viel verbraucht, wo sich das stärkere Modell lohnt und wo das sparsamere genügt – und daraus bauen wir schrittweise unsere Best Practices für die KI-Nutzung auf.

Wie geht es weiter?

Das Problem und die möglichen Lösungen entfalten wir in den kommenden Wochen Rolle für Rolle in eigenen Artikeln.

Fortsetzung folgt.

Die wichtigsten Themen, so weit sie sich heute abzeichnen

  • Technik und Delivery — Wie sich der Tokenverbrauch messbar machen lässt und welche Schlüsse sich daraus ziehen lassen.
  • Vertrieb und Marktrealität — Was wir am Markt sehen: die Verschiebung hin zu verbrauchsbasierten Modellen und was daraus in Kombination mit den sinkenden Tokenpreisen folgt.
  • Unternehmerische Perspektive — Wo sich das alles finanziell rechnet und wo die realen Schmerzpunkte auf CEO-Ebene liegen.
  • Bewusste KI-Nutzung — Welches Werkzeug und welches Modell sich für welche Aufgabe eignet – einschließlich unserer Erfahrungen mit dem eigenen Hermes-Agenten.
  • Onboarding und HR — Wie wir Kolleginnen und Kollegen in der Entwicklung und darüber hinaus zur richtigen, effizienten KI-Nutzung befähigen.

Newsletter
Get new articles delivered to your inbox.
A concise monthly brief: the latest articles, audit insights, and event invitations. Unsubscribe with one click—no spam.
Thank you! Your submission has been successfully recorded!
Oops! Something went wrong while submitting the form.