Insights

AI cost is consumption, not licences: why 93% of organizations blow the budget

AI is the first line in your technology budget that grows with usage instead of headcount. Every piece of software before it charged per seat and was predictable; this one charges per action and has no ceiling of its own. That is why 93% of organizations exceed their AI budget, and why the problem shows up exactly when the project starts working.

7 min read
A cost chart over time: a flat stepped line for per-seat software licences stays under the budget ceiling, while an AI consumption curve climbs sharply, crosses that ceiling and keeps rising

The moment AI cost becomes a problem is not when the project fails. It is when it works. The pilot was cheap because twenty people used it on selected cases; now fifteen hundred people use it every day, and last month's invoice arrived with a number nobody on the committee had seen before. The project is succeeding, and that is why it is at risk.

The short answer: AI is the first line in the technology budget that grows with usage instead of headcount. Everything before it charged per seat, so the annual cost was known before signing. This one charges per action, so usage writes the invoice. Managing it is not squeezing vendors, it is treating consumption the way any strategic asset is treated.

Why does cost explode exactly when the project works?

Because scaling does not multiply the price, it multiplies the consumption.

Moving from isolated use cases to enterprise-wide adoption, AI spend increases nearly fourfold [1]. Not because the vendor raised rates, but because the number of calls multiplies: more users, more times per day, with longer conversations and agents making several calls per task.

The consequence is measured, and it is uncomfortable.

FIG-1Nearly everyone blows the budget, one in four has the practice to avoid itOrganizations, %
Data
Value
Exceed their AI budget93%
Have mature AI FinOps practices23%

Ninety-three against twenty-three. The gap between the two figures is the definition of the problem: almost everyone has the spend, almost nobody has the discipline to manage it.

And it is not a passing spike. 62% of organizations have already moved beyond experimentation into active deployment, and most expect their AI spend to rise at least another 25% over the next twelve months [1].

What makes this cost different from any other software?

Its invoice is set by usage, not by your purchasing decision. That is a first for a technology budget.

Traditional software AI
Billing unit Seat or subscription Token or action
What makes it grow Hiring people Using the system more
When the cost is known Before signing At month close
Who can trigger it Procurement Any user, or an agent in a loop
Natural ceiling The size of the payroll None

The last row deserves reading twice. A badly designed agent that repeats a step does not produce a visible error: it produces an invoice. And it produces it while every dashboard stays green.

Where is the spend nobody sees?

Spread across vendors, and it is more than it looks.

Between 20 and 30% of AI spend typically goes unaccounted for, scattered across fragmented vendors [1]. This is not fraud or serious negligence: it is what happens when four teams sign up with four different vendors on four different cards, and nobody consolidates.

It is also the cheapest saving available, because it requires optimizing nothing. It requires looking.

How much comes back, and from which lever?

A meaningful amount, and the figures come as ranges because they depend on the starting point.

FIG-2What each lever gives back, observed rangeReduction on spend, %
Data
LeverMinimumMaximumDelta
Thoughtfully managed consumption20%30%10%
Unallocated spend surfaced by consolidating20%30%10%
Sourcing and negotiation (unit cost)10%20%10%

On top of those sits a technical lever that deserves its own name: prompt caching can cut input costs by roughly 90% [1]. If your agent sends the same context block on every call, and almost all of them do, that single configuration changes the order of magnitude of the bill.

The remaining spend-optimization levers are model selection per task, token controls, agent management, batching, and open-weight models where they apply [1].

What do you do first without a FinOps team?

Three things that fit in a day, and they capture most of the saving.

Worth saying plainly: the full five-lever framework is written for an organization with a CIO, a cross-functional team and formal governance. Most companies do not have that, and telling them to stand up a permanent FinOps capability before touching anything is a way of doing nothing.

The minimal version, in order:

  1. Consolidate in one place how much you spend and on what. A spreadsheet works. What does not work is the number living on four credit cards.
  2. Set a spend cap per use case, with an alert before the ceiling. That is the difference between finding out at month close and finding out on Tuesday.
  3. Turn on prompt caching. It is configuration, not architecture, and it has the best effort-to-return ratio of any lever here.

Then, and only then, comes the question that saves the most over time: does each task still need the large model? Almost never. Reviewing that monthly is worth more than an annual negotiation exercise.

Why does this weigh more in Latin America?

Because the margin for error is smaller and the conversion from investment to impact is already tight.

In the region, about one in ten firms uses artificial intelligence and only 6% captures significant value from it [2]. When the cost of operating quadruples at scale, that distance between using and capturing value does not hold: it becomes the reason the project gets cancelled before reaching the scale where it would have paid for itself.

Put differently: here, cost discipline is not a later optimization. It is part of whether the project arrives at all.

How we apply this at MasterDragon

We budget per action, not per project.

In practice, before writing code we estimate what one complete run of the flow costs: how many calls it makes, with which model, with how much context. That number times the expected volume is the real budget, and it often changes the design before we start, because it reveals that an expensive step can be handled by a small model or by no model at all.

We also instrument cost from the first deployment, next to the traces. A system that does not report what each run costs is a system that will surprise you, and the surprise always lands in the month the project finally got used for real.

If you are about to scale an AI use case and want to know what it will cost before the invoice tells you, talk to our engineers. You can start with how we build your software and review our portfolio of shipped AI-native products. Model routing per task, the largest technical lever, is one of the layers in the 13 layers of a real AI system; what to change in operations so the investment shows up in the result is in the five operating-model changes; and why to build rather than rent when the process is yours, in custom software is an asset, not an expense.

References

  1. Sachdeva, P., Lala, W., Javaji, A., Takkar, K., & Arora, P. (2026, July 20). The cost of intelligence: How CIOs can manage AI demand at scale. McKinsey & Company. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-cost-of-intelligence-how-cios-can-manage-ai-demand-at-scale
  2. Center for Latin America Convergence. (2026, July). Unlocking the Digital Potential of Latin America and the Caribbean: Market Integration, Investment, Smart Regulation and Artificial Intelligence.

Frequently asked questions

Why does AI cost explode at scale?

Because it is billed on consumption, not seats. Moving from isolated use cases to enterprise adoption, AI spend increases nearly fourfold, and 93% of organizations end up exceeding their AI budget. The curve grows with usage, so the moment of greatest financial risk is exactly when the project starts working well.

What makes AI cost different from any other software cost?

Its bill is set by usage, not by your purchasing decision. Traditional software charges per seat or subscription: you know the annual cost before signing and it grows with headcount. AI charges per token or per action, so usage writes the invoice and there is no natural ceiling. It is the first technology budget line that can grow without anyone making a buying decision.

How much AI spend goes unaccounted for?

Between 20 and 30% typically goes unallocated, scattered across fragmented vendors. And only 20 to 25% of companies have mature AI FinOps practices, meaning the discipline required to see and control it.

How much can optimization recover?

Between 20 and 30% through thoughtful consumption management, and about a third of surveyed organizations have already achieved that through active optimization. Sourcing and negotiation add a further 10 to 20% reduction in unit cost. Prompt caching alone can cut input costs by roughly 90%.

What is AI FinOps?

The practice of managing AI consumption the way any strategic asset is managed: consolidated spend visibility, demand forecasting, continuous optimization and assigned accountability per area. It is not a monthly report, it is a permanent cross-functional team.

What does a small company without a FinOps team do?

The same as large ones, in minimal form: consolidate in one place how much is spent and on what, set a spend cap per use case, turn on prompt caching, and review monthly whether the large model is still needed for each task. The first three take a day and capture most of the savings.

About the author

MasterDragon Engineering Team

MasterDragon Engineering Team

AI Engineering Team · MasterDragon.AI

The MasterDragon Engineering Team designs and ships production-grade agentic AI systems for companies in LATAM and the US: custom AI-native software, WhatsApp agents, internal copilots and end-to-end operations automation, with measurable reliability and KPIs.