A Glosternél, és gondolom Nálatok is a napi fejlesztői és üzleti munkába egyre mélyebben épül be az AI, jellemzően Claude Code- és Microsoft Copilot-alapú, emberi felügyelet melletti, vagy már néha sokáig önállóan futó, ún. “agentic” munkafolyamatokban. Ahogy nő a használat, egyre nehezebb megkerülni a kérdést: mennyibe kerül mindez, és hogyan tarthatóak kézben a költségek? És főleg, hogyan használjuk, hogy megérje? A probléma már itt van, egyre fájdalmasabb. Egy belső kick-off beszélgetésen több szempontból körül is jártuk: hogyan gazdálkodjunk az AI-tokenhasználattal úgy, hogy a napi munkát ne akasszuk meg, az AI használat költsége viszont ne szaladjon el? Jó kis vita kerekedett belőle, de a végén mégis egy közös álláspontra jutottunk.
Ahhoz, hogy a kérdést reálisan lássuk, előbb a piaci irányt érdemes megnézni.
A trend a fogyasztásarányos modell felé mutat. Az elmúlt hónapokban az a kép rajzolódott ki, hogy a nagy szállítók egyre inkább usage alapú árazásra állnak át, és a Microsoft Copilot esetében ez már ma is konkrétan látszik. Ha ez az irány folytatódik, minden szállító és minden modelljére kiterjed, előbb-utóbb mindenkinek a pay-as-you-go modellel kell számolnia. Érdemes tehát nemcsak az első számla megérkezése után pánikolni, hanem előre felkészülni rá.
Nem mindegy, mire használjuk az AI modelleket. Az egyszeri kérdés-válasz típusú munka olcsó. Az agent-alapú folyamatok viszont, ahol a modell több lépésben, önállóan dolgozik, nagyságrendekkel több tokent fogyasztanak: minden iteráció, minden beolvasott fájl, minden újratervezés terheli a keretet. Ahogy a napi munka egyre nagyobb részét visszük ilyen folyamatokba, a fogyasztás nem lineárisan, hanem ugrásszerűen nő. Éppen ezért nem elég annyit tudni, hogy „használunk AI-t”. Az számít leginkább, hogy mire, mennyit, és milyen hatékonysággal használjuk.
A fejlesztők idő előtti bekorlátozása kontraproduktív. A technikai csapat álláspontja szerint először meg kell tanulni jól használni az eszközöket, és csak utána szabad költséget optimalizálni. A szoftverfejlesztésben régi igazság, hogy az idő előtti optimalizálás a projekt halála, és ez itt is áll. Ha túl korán húzunk szűk kereteket, épp azt a tanulási fázist fojtjuk el, amikor a csapat még csak most találja meg, hol adja az eszköz a legtöbbet. A költségoptimalizálásba akkor érdemes belevágni, amikor a szervezet AI-érettsége már megvan hozzá.
A pay-as-you-go árazási modell rugalmas, de a vevők számára kevésbé kiszámítható. Ebben a modellben jellemzően workspace szinten lehet korlátokat állítani, cserébe a fogyasztás nehezebben tervezhető, és a névleges korlátok mellett is érhetnek meglepetések a számlázásban. Ez önmagában is azt támasztja alá, hogy a tényleges használatot tervezni, monitorozni, kontrollálni kell, különben csak a számla megérkezésekor derül ki, mennyi az annyi.
Az előre megvásárolt keret kiszámítható, de ha idő előtt elfogy, a napi munkát könnyen megakaszthatja. A licencalapú konstrukciók, például a team license, előnye, hogy a költségkeret előre rögzített, és sok csapatnak és az idő nagy részében ez elég is. De ha megkérdezed a sales vezetődet, tuti azt mondja, hogy a napi tokenkeret a legrosszabb pillanatban, pl. egy sürgős ajánlatadás utolsó órájában fog elfogyni. És mire a keret újratöltődik, addigra lejár a határidő, kiesel a tenderből, vagy az ügyfél oldalán már rég tárgytalan a kérdés. És visszautalva a piaci trendre, az is nyitott, meddig maradnak egyáltalán ezek a konstrukciók.
Talán nem kell mindenkinek minden modell. A technikai csapaton belül felmerült, hogy egy szervezeten belül nem biztos, hogy minden területnek a legerősebb modellre van szüksége. A pénzügynek például elég lehet egy takarékosabb modell, és ezzel a hozzáférés eleve, természetes módon szabályozódik, anélkül hogy külön korlátokat kellene bevezetni. A wrapper-megoldásoknak nem vagyunk hívei, de a modellréteg szerinti differenciálás reális irány: a feladathoz igazított modellválasztás egyszerre csökkenti a költséget és tereli a használatot a józan mederbe.
A beszélgetés végére egy dologban egyetértettünk. Bármelyik irányba is indulunk el, pay-as-you-go, licencalapú keret vagy modellréteg szerinti differenciálás, mérés nélkül vakon tapogatózunk.
A mérés mutatja meg, ki adaptálódott ténylegesen, mely területen és milyen feladatokra megy el a fogyasztás, hol van pazarlás és hol reális a megtérülés. Enélkül minden korlátozás és annak elvetése is csak megérzésen alapul, sima lottó, rulett, csapdlecsacsi. Amit nem mérünk, azt nem is tudjuk kézben tartani.
A Glosternél így már nem is elméleti kérdésként tekintünk erre: elkezdtük mérni és optimalizálni a saját AI-használatunkat. Monitorozzuk, elmentjük és elemezzük, melyik csapat mire és mennyit fogyaszt, hol éri meg az erősebb modell, és hol elég a takarékosabb, és ebből fokozatosan építjük ki a legjobb AI használati gyakorlatot. A célunk egyrészt, hogy spóroljunk egy kalap pénzt (mert ahogy írtuk, szeretjük...), másrészt, hogy irányt mutassunk, és menet közben megosszuk, mi az, ami nálunk működik, és mit javaslunk, hogy más is kipróbáljon, bevezessen. Mert ez a probléma Téged is érint, és mindenki mást is, aki a cége hatékonyságát az AI-ra építi. Szóval kb. mindenkit.
Akkor most mi lesz? A felvetett problémát és a lehetséges megoldásokat a következő hetekben szerepenként, külön cikkekben bontjuk ki. A főbb témák, illetve ami már körvonalazódik:
Folytatjuk!
At Gloster, and I imagine at your company too, AI is being built ever more deeply into everyday development and business work, typically in Claude Code and Microsoft Copilot based workflows: sometimes under human supervision, sometimes already running on their own for long stretches, in what are now called "agentic" setups.
As usage grows, one question becomes harder and harder to avoid: how much does all this cost, and how can those costs be kept under control? And above all, how do we use it so that it actually pays off?
The problem is already here, and it is getting more painful. We looked at it from several angles in an internal kick-off discussion: how do we manage AI token usage in a way that does not hold up the daily work, but also does not let the cost of using AI run away from us? It sparked a good debate, but in the end we did arrive at a shared position.
Before we argue about control, it is worth reading the market.
The direction of travel is consumption-based. Over the past few months the picture has sharpened: the large suppliers are moving towards usage-based pricing, and with Microsoft Copilot you can already see it in the open. If that continues, spreading across every supplier and every model, pay-as-you-go becomes the default everyone plans around. So the smart move is to prepare for it now, rather than panic when the first bill lands.
What you use AI models for makes all the difference. A single question-and-answer exchange is cheap. Agent-based workflows — where the model works across many steps on its own — burn tokens on another order entirely: every iteration, every file it reads, every replan draws down the budget. As more of the daily work shifts into those workflows, consumption does not climb in a straight line. It jumps.
Knowing that "we use AI" tells you almost nothing. What counts is what you use it for, how much, and how well.
Fencing developers in too early is counterproductive. Software engineering has an old truth that premature optimisation kills the project, and it holds here too. Draw the budget tight too soon and you smother the exact learning phase where the team is still finding where the tool earns its keep. Cost optimisation is worth starting once the organisation has the AI maturity to support it.
Pay-as-you-go is flexible, but harder for the buyer to predict. In this model you can usually set limits at workspace level, and in return consumption is tougher to forecast. Even with nominal limits in place, the billing can still surprise you. That alone makes the case for planning, monitoring and controlling actual usage.
A pre-purchased allowance is predictable, but it can stall the day's work if it runs out early. Licence-based arrangements fix the budget up front, and for many teams, most of the time, that is enough. But the daily token allowance can run dry at the worst possible moment — in the final hour before an urgent bid goes out. And it is an open question how long these arrangements survive the market shift at all.
Perhaps not everyone needs every model. Not every function in an organisation needs the strongest model. Finance, for instance, may be well served by a leaner one, and that choice regulates access naturally, without imposing a separate cap. Matching the model to the task cuts the cost and steers usage back into sane territory at the same time.
By the end of the conversation we agreed on one point. Whichever way we go — pay-as-you-go, a licence-based allowance, or differentiation by model tier — without measurement we are feeling our way in the dark.
Measurement shows who has genuinely adopted, which areas and which tasks the consumption goes on, where there is waste and where the return is real. Without it, every limit — and every decision to lift one — rests on nothing but a hunch.
What you do not measure, you cannot keep in hand.
So at Gloster we no longer treat this as theory. We have started measuring and optimising our own AI usage. We monitor, log and analyse which team spends what and on what, where the stronger model pays for itself and where the leaner one is enough, and from that we are building our best-practice AI playbook step by step.
Our aim is partly to save a serious pile of money, and partly to point the way, and to share as we go what works for us and what we would suggest others try out and adopt. Because this problem affects you too, and everyone else who is building their company's efficiency on AI. So, pretty much everyone.
Over the upcoming weeks we will unpack the problem and the possible solutions role by role, in separate articles.
To be continued!
Bei Gloster, und vermutlich bei Ihnen auch, wächst KI immer tiefer in den Entwicklungsalltag und in die Arbeit der Fachbereiche hinein. Typischerweise in Workflows auf Basis von Claude Code und Microsoft Copilot, unter menschlicher Aufsicht – manchmal aber auch schon in agentischen Workflows, die über längere Zeit selbstständig laufen.
Je stärker die Nutzung steigt, desto schwerer lässt sich die Frage umgehen: Was kostet das alles, und wie behält man die Kosten im Griff?
Das Problem ist längst da, und es tut zunehmend weh. In einem internen Kick-off haben wir es aus mehreren Blickwinkeln beleuchtet: Wie gehen wir mit dem Tokenverbrauch um, ohne die tägliche Arbeit auszubremsen und ohne dass die Kosten davonlaufen? Daraus wurde eine ordentliche Debatte – am Ende aber auch ein gemeinsamer Standpunkt.
Der Trend geht in Richtung verbrauchsabhängiger Modelle. In den letzten Monaten hat sich abgezeichnet, dass die großen Anbieter zunehmend auf verbrauchsbasierte Preise umstellen; bei Microsoft Copilot ist das heute schon konkret sichtbar. Setzt sich diese Richtung fort und erfasst sie jeden Anbieter und jedes Modell, muss früher oder später jeder mit Pay-as-you-go rechnen.
Es lohnt sich also, nicht erst nach der ersten Rechnung in Panik zu geraten, sondern sich vorab darauf einzustellen.
Es macht einen Unterschied, wofür wir die Modelle einsetzen. Einzelne Frage-Antwort-Interaktionen sind günstig. Agentische Abläufe dagegen, bei denen das Modell in mehreren Schritten selbstständig arbeitet, verbrauchen um Größenordnungen mehr Token: jede Iteration, jede eingelesene Datei, jede Neuplanung belastet das Budget. Je größer der Anteil solcher Abläufe an der täglichen Arbeit wird, desto stärker steigt der Verbrauch – nicht linear, sondern sprunghaft.
Deshalb genügt es nicht zu wissen, dass wir KI nutzen. Entscheidend ist, wofür, wie viel und mit welcher Effizienz.
Entwickler zu früh einzuschränken ist kontraproduktiv.Zuerst muss man lernen, die Werkzeuge gut zu nutzen, erst danach optimiert man die Kosten. Vorzeitige Optimierung erstickt genau die Lernphase, in der das Team herausfindet, wo das Werkzeug den größten Nutzen bringt.
Pay-as-you-go ist flexibel, für Kunden aber weniger planbar.Limits lassen sich auf Workspace-Ebene setzen, doch der Verbrauch ist schwerer zu planen, und selbst bei nominellen Grenzen können in der Abrechnung Überraschungen auftreten.
Ein vorab gekauftes Kontingent ist planbar – geht es zu früh zur Neige, blockiert es aber die tägliche Arbeit.Das tägliche Tokenkontingent geht im schlechtesten Moment zur Neige, etwa in der letzten Stunde vor einer dringenden Angebotsabgabe.
Vielleicht braucht nicht jeder jedes Modell.Innerhalb einer Organisation benötigt nicht jeder Bereich das stärkste Modell. Die Wahl des passenden Modells je Aufgabe senkt die Kosten und lenkt die Nutzung zugleich in vernünftige Bahnen.
Am Ende des Gesprächs waren wir uns in einem Punkt einig. In welche Richtung wir auch gehen – Pay-as-you-go, lizenzbasiertes Kontingent oder Differenzierung auf Modellebene –, ohne Messung tappen wir im Dunkeln.
Die Messung zeigt, wer die Werkzeuge tatsächlich nutzt, in welchem Bereich und für welche Aufgaben der Verbrauch anfällt, wo Verschwendung entsteht und wo sich der Einsatz realistisch rechnet.
Was wir nicht messen, können wir auch nicht im Griff behalten.
Bei Gloster betrachten wir das deshalb nicht mehr als theoretische Frage: Wir haben begonnen, unsere eigene KI-Nutzung zu messen und zu optimieren. Wir erfassen, speichern und analysieren, welches Team wofür und wie viel verbraucht, wo sich das stärkere Modell lohnt und wo das sparsamere genügt – und daraus bauen wir schrittweise unsere Best Practices für die KI-Nutzung auf.
Das Problem und die möglichen Lösungen entfalten wir in den kommenden Wochen Rolle für Rolle in eigenen Artikeln.
Fortsetzung folgt.