Brief
|
|
概要
In part I, we laid out the provocation that the opex mix is shifting from headcount to tokens and that there is no clear path of transition. In software engineering, the domain furthest along in AI enablement, spending on tokens is only about 1% to 2% the cost of headcount. The pattern is the same across sales, support, strategy, and ops: at 1% and talking about getting to 20% to 30% (see Figure 1). But where we are looks radically different depending on which industry you’re in and where you sit within the organization. Three levels, three different problems, one transformation.
Figure 1
At the CEO level, the destination is clear. Intuit’s CEO has set a vision to become a builder powerhouse by tripling developer productivity. Across our tech-forward client base, the mindset has flipped from “you helping AI” to “AI helping you.” Alpha teams redesign the workflow first, prove the velocity gain, then expand to the broader organization. These aren’t experiments; they’re mandates. And the leaders setting them aren’t debating the returns on investment. They’re impatient with the pace of organizational change beneath them. At the general manager (GM) level, it’s a budget-and-speed problem. Where do I find the unbudgeted millions this quarter? Token spending doesn’t fit an existing line item and requires an approval chain that doesn’t exist. The pilot was stellar, but scaling to 10,000 seats means navigating procurement, security, legal, data, and IT cycles that were built for an annual cadence. The GM knows the destination, but the path is littered with quarterly budget reviews, vendor approvals, and organizational redesigns that can’t happen fast enough. At the individual level, there’s a chasm between the heroes and everyone else. Early data suggests that the top 5% of users in each company often consume more tokens than the other 95% combined. In some cases, senior staff engineers, chief architects, sales rainmakers, and strategy directors say, “I don’t need a team; they only slow me down.” It’s unclear how we get everyone to cross the chasm or how to manage superusers’ costs without slowing them down.
The cost-per-task paradoxHere’s what makes this truly uncomfortable: In some domains, agent and token costs are already more expensive than offshore human resources. Not everywhere: These are still sparse domains. But they exist. Agents consume significant tokens on multistep reasoning, error correction, and context loading, which add up fast on complex workflows. The speed and quality gains are real, but the per-task economics don’t always pencil out, particularly when done for the wrong tasks or without the appropriate orchestration. The question is where you’re paying a premium for capability vs. where you should be getting a cost arbitrage. So, will the cost come down? The hope is obvious: Model prices are falling roughly 10 times per generation. But here’s the reality of what we’re actually seeing: The effective cost per task is often staying flat. Why? Three forces are working against the headline price declines. First, everyone stays on frontier. When the next Claude or GPT ships, nobody says, “Great, I’ll keep using the old model and pocket the savings.” They upgrade. Last generation gets cheaper; frontier stays expensive (see Figure 2). Second, tokens per query keep climbing as agents take on more complex, multistep work, such as orchestrating tool calls, correcting errors, and loading context, with more advanced models consuming more tokens on more difficult problems. Third, usage expands: Once a team discovers what agents can do, they find 10 more workflows to throw at them.
Figure 2
Note: MMLU-Pro is massive multitask language understanding, a benchmark used to evaluate large language models 出所 Bain analysisThe trend is more nuanced than the headline. The cost of tokens (typically measured in millions) fell by half from December 2024 to December 2025 while tokens consumed grew by 4.5 times over the same period (see Figure 3).
Figure 3
Net-net: The models get less expensive per token, the usage gets heavier per task, and the bill stays stubbornly high. The hope is that this is like 2G/3G costs circa 2009: expensive now, destined to plummet. And maybe it is. But every six months feels like it should be the inflection point. It never quite is.
What would swing itThe 70/30 hypothesis, reflecting the cost of headcount and cost of tokens, is a scenario, not a forecast. Whether it materializes and how fast both depend on a handful of variables that interact in nonlinear ways. The range of outcomes is enormous. AT&T learned this firsthand: At 8 billion tokens a day, the company said it reorchestrated so that large "super agents" route tasks to smaller, domain-specific worker models instead of pushing everything through frontier. The company reported a 90% cost reduction and three times throughput, not by using less AI but by right-fitting the model to the job. One architectural decision moved the economics by an order of magnitude. That's the kind of nonlinearity we're talking about. Here are some of the variables we watch most closely.
Navigating the token economicsUncertainty is real; it’s not an excuse for inaction. Here are five moves that hold up across scenarios.
The bottom lineThe opex shift from headcount to tokens isn’t a budget problem; it’s a structural transformation. The economics are unsettled, and the path is nonlinear. That's not a reason to wait. It's the reason to start instrumenting now so that when the curve breaks, you're navigating with data instead of intuition. Next Monday: Pull your top 10 SaaS contracts and your token spend to date. Instrument one workflow end to end: What does it actually cost per task, per outcome? That’s the number nobody in your org knows yet. Once you have it, every other decision gets easier. Explore this series |