Technology Report
|
|
Резюме
This article is part of Bain’s Technology Report 2026 The first time a company receives a startling AI bill, the instinct is almost always the same: Put limits on usage. Many companies set up per-person budgets and implement approval processes to clamp down on costs before they spiral out of control. The reaction seems to make sense, but it solves the wrong problem. Companies can become so concerned about controlling costs that they restrict the very behaviors they’re trying to encourage. Developers ration AI use, sales teams stop experimenting, and leaders worry more about the bill than the value. Our view is that the bigger problem right now in most organizations isn’t runaway token spending, but too much spending in the wrong places. Most companies should be spending tokens across more of their organization to radically change their business. Across organizations, underuse is the bigger problem. One global technology company set a soft cap of $500 per month for each developer to help encourage AI use; the average cost landed around $200. At another tech company, engineering leaders were metered so cautiously that they left their allocated $40-per-person monthly budget unspent, never touching the buffer pool sitting behind it. Leadership eventually told them to stop metering and go spend. But this restraint is often overshadowed by a few big spenders within a small number of AI workflows that are not spread evenly across thousands of users. A lack of transparency limits the ability to measure returns and channel spending to the places where it will generate the most value. For one widely deployed knowledge-assistant tool, nearly all tokens are cached input tokens (embedded and reused in API requests) that a company can’t easily see or forecast. Managing token spending effectively requires precision, not blanket controls. Manage workflows, not usersThe right question isn’t "how much should each employee spend?" It’s "which workflows deserve more AI, and which deserve less?" Framing the question in such a way leads to three distinct approaches, each with its proper place in the company (see Figure 1).
Figure 1
The principle is simple: Govern the risky tail tightly, and get out of the way almost everywhere else. Measure cost by workflowThis approach only works if organizations understand what each workflow actually costs and what it’s worth. Most companies track AI spending at the user or business unit level, but that doesn’t show where value is being created. The more useful question is “what does it cost to complete a proposal, resolve a customer issue, review a contract, or generate production-ready code?” First versions of agentic workflows often cost far more than expected because prompts, orchestration, and model choices haven’t been optimized yet. In our experience, a new workflow might cost 10 times more than it ultimately should. Once teams understand the economics, they can often remove most of that excess through better architecture, smarter model selection, and more deterministic workflow design. As token prices declined, companies expected their costs to follow. But most companies typically upgrade to the latest frontier models, agents consume more tokens as they tackle increasingly sophisticated tasks, and adoption expands into more workflows. As a result, lower per-token costs are offset by higher usage, leaving overall cost per task stubbornly high (see Figure 2). The expectation is that these economics will eventually improve, but the anticipated cost inflection point has yet to materialize (for more, read the Bain Brief “How Token Economics Will Change Opex”).
Figure 2
Even so, the right strategy isn’t just to use cheaper models; it’s to reserve expensive reasoning for the parts of a workflow that genuinely require it while making the rest more structured and predictable. Govern selectivelyThis also changes how organizations should govern AI spending. Most workflows belong with the business. Engineering leaders should manage engineering workflows, and sales leaders should manage sales workflows. They’re best positioned to understand where AI creates value, and they can encourage adoption without unnecessary bureaucracy (for more, read the Bain Brief “Mobilizing the Organization in the Agent Economy”). Only the highest-risk workflows require central oversight. A small governance team can establish standards for testing, monitoring, and cost controls while ensuring that open-ended agentic systems include appropriate safeguards before they are deployed broadly. Where possible, those controls should be built into the core AI platforms themselves, making safe, cost-effective behavior the default rather than relying on users to follow policy. The same principle applies to spending limits. Low caps often create exactly the wrong behavior, discouraging experimentation and slowing adoption. High soft limits, combined with clear visibility into unusual consumption, encourage teams to use AI where it creates value while still protecting against genuinely excessive costs. Prepare for the next waveMost organizations are still in the early stages of AI adoption. As autonomous agents become more capable, AI consumption is likely to increase dramatically. Persistent agents, long-running optimization tasks, and increasingly sophisticated workflows will all drive significantly higher token usage than today's interactive assistants. That makes it even more important to establish the right management approach now. Organizations that rely on blunt cost controls will struggle to capture AI's full value. Those that understand the economics of their workflows will be able to expand adoption confidently while managing the relatively small number of use cases in which costs can escalate quickly. More from the report
Read our Technology Report 2026 |