The moment AI cost becomes a problem is not when the project fails. It is when it works. The pilot was cheap because twenty people used it on selected cases; now fifteen hundred people use it every day, and last month's invoice arrived with a number nobody on the committee had seen before. The project is succeeding, and that is why it is at risk.
The short answer: AI is the first line in the technology budget that grows with usage instead of headcount. Everything before it charged per seat, so the annual cost was known before signing. This one charges per action, so usage writes the invoice. Managing it is not squeezing vendors, it is treating consumption the way any strategic asset is treated.
Why does cost explode exactly when the project works?
Because scaling does not multiply the price, it multiplies the consumption.
Moving from isolated use cases to enterprise-wide adoption, AI spend increases nearly fourfold [1]. Not because the vendor raised rates, but because the number of calls multiplies: more users, more times per day, with longer conversations and agents making several calls per task.
The consequence is measured, and it is uncomfortable.
Data
| Value | |
|---|---|
| Exceed their AI budget | 93% |
| Have mature AI FinOps practices | 23% |
Ninety-three against twenty-three. The gap between the two figures is the definition of the problem: almost everyone has the spend, almost nobody has the discipline to manage it.
And it is not a passing spike. 62% of organizations have already moved beyond experimentation into active deployment, and most expect their AI spend to rise at least another 25% over the next twelve months [1].
What makes this cost different from any other software?
Its invoice is set by usage, not by your purchasing decision. That is a first for a technology budget.
| Traditional software | AI | |
|---|---|---|
| Billing unit | Seat or subscription | Token or action |
| What makes it grow | Hiring people | Using the system more |
| When the cost is known | Before signing | At month close |
| Who can trigger it | Procurement | Any user, or an agent in a loop |
| Natural ceiling | The size of the payroll | None |
The last row deserves reading twice. A badly designed agent that repeats a step does not produce a visible error: it produces an invoice. And it produces it while every dashboard stays green.
Where is the spend nobody sees?
Spread across vendors, and it is more than it looks.
Between 20 and 30% of AI spend typically goes unaccounted for, scattered across fragmented vendors [1]. This is not fraud or serious negligence: it is what happens when four teams sign up with four different vendors on four different cards, and nobody consolidates.
It is also the cheapest saving available, because it requires optimizing nothing. It requires looking.
How much comes back, and from which lever?
A meaningful amount, and the figures come as ranges because they depend on the starting point.
Data
| Lever | Minimum | Maximum | Delta |
|---|---|---|---|
| Thoughtfully managed consumption | 20% | 30% | 10% |
| Unallocated spend surfaced by consolidating | 20% | 30% | 10% |
| Sourcing and negotiation (unit cost) | 10% | 20% | 10% |
On top of those sits a technical lever that deserves its own name: prompt caching can cut input costs by roughly 90% [1]. If your agent sends the same context block on every call, and almost all of them do, that single configuration changes the order of magnitude of the bill.
The remaining spend-optimization levers are model selection per task, token controls, agent management, batching, and open-weight models where they apply [1].
What do you do first without a FinOps team?
Three things that fit in a day, and they capture most of the saving.
Worth saying plainly: the full five-lever framework is written for an organization with a CIO, a cross-functional team and formal governance. Most companies do not have that, and telling them to stand up a permanent FinOps capability before touching anything is a way of doing nothing.
The minimal version, in order:
- Consolidate in one place how much you spend and on what. A spreadsheet works. What does not work is the number living on four credit cards.
- Set a spend cap per use case, with an alert before the ceiling. That is the difference between finding out at month close and finding out on Tuesday.
- Turn on prompt caching. It is configuration, not architecture, and it has the best effort-to-return ratio of any lever here.
Then, and only then, comes the question that saves the most over time: does each task still need the large model? Almost never. Reviewing that monthly is worth more than an annual negotiation exercise.
Why does this weigh more in Latin America?
Because the margin for error is smaller and the conversion from investment to impact is already tight.
In the region, about one in ten firms uses artificial intelligence and only 6% captures significant value from it [2]. When the cost of operating quadruples at scale, that distance between using and capturing value does not hold: it becomes the reason the project gets cancelled before reaching the scale where it would have paid for itself.
Put differently: here, cost discipline is not a later optimization. It is part of whether the project arrives at all.
How we apply this at MasterDragon
We budget per action, not per project.
In practice, before writing code we estimate what one complete run of the flow costs: how many calls it makes, with which model, with how much context. That number times the expected volume is the real budget, and it often changes the design before we start, because it reveals that an expensive step can be handled by a small model or by no model at all.
We also instrument cost from the first deployment, next to the traces. A system that does not report what each run costs is a system that will surprise you, and the surprise always lands in the month the project finally got used for real.
If you are about to scale an AI use case and want to know what it will cost before the invoice tells you, talk to our engineers. You can start with how we build your software and review our portfolio of shipped AI-native products. Model routing per task, the largest technical lever, is one of the layers in the 13 layers of a real AI system; what to change in operations so the investment shows up in the result is in the five operating-model changes; and why to build rather than rent when the process is yours, in custom software is an asset, not an expense.
References
- Sachdeva, P., Lala, W., Javaji, A., Takkar, K., & Arora, P. (2026, July 20). The cost of intelligence: How CIOs can manage AI demand at scale. McKinsey & Company. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-cost-of-intelligence-how-cios-can-manage-ai-demand-at-scale
- Center for Latin America Convergence. (2026, July). Unlocking the Digital Potential of Latin America and the Caribbean: Market Integration, Investment, Smart Regulation and Artificial Intelligence.

