BIP America News & Media Platform

collapse
Home / Daily News Analysis / Why AI tokens will send your enterprise cloud bill sky-high again

Why AI tokens will send your enterprise cloud bill sky-high again

Jun 28, 2026  Twila Rosenbaum  24 views
Why AI tokens will send your enterprise cloud bill sky-high again

AI tokens will remind many enterprise customers of cloud pricing's early days. However, measuring the value derived from AI remains an unsolved problem. SAN DIEGO -- A few months ago, most people paid a flat fee for their AI access. That was then. This is now. The days of AI pricing as a loss-leader are over. As everyone has discussed here at FinOps X 2026, AI's token-based pricing model is becoming the foundation of the entire generative AI economy, and it's far more expensive than older models.

Key facts about AI token pricing

  • Token as the atomic unit: J.R. Storment, executive director of the FinOps Foundation, calls tokens "the atomic unit of AI." A token is the smallest unit a word or phrase can be broken down into when processed by an LLM. For English, 100 tokens ≈ 75 words.
  • Pricing models shift: Labs charge per million input tokens and per million output tokens, hiding complexity in model choice, quantization, caching, and agent usage. The all-you-can-eat token era is over.
  • Cost explosion: Global token usage grew linearly until late 2025, then exploded with larger context windows (millions of tokens per conversation) and agentic patterns (loops, retries, corrections). Some $200/month power users cost tens of thousands of dollars monthly.
  • Token price trends: Since 2023, token prices have fallen dramatically. However, since November 2025, prices have flattened due to hardware supply constraints (GPUs, power) that may not ease until 2028.
  • Jevons paradox: Falling unit costs drive exploding total spend. Goldman Sachs estimates global token use rising from 6 quadrillion today to 120 quadrillion within 3.5 years.
  • FinOps challenges: AI does not just stretch the cloud playbook, it breaks it. SAP built a custom AI FinOps framework with three pillars: spend visibility, economics (token-level metrics like input/output ratios, cached token ratios, token-to-spend drift), and value (cost per use case, inference cost by revenue).
  • Tokenomics beyond FinOps: The Linux Foundation is forming a Tokenomics Foundation to normalize measurement and allocation. Tokenomics covers production (energy+capital → tokens), consumption (allocation, forecasting, optimization), and value (monetization, pricing, labor implications).
  • Business model abstraction: Vendors layer credits, hybrid subscription+usage, or direct pass-through models on top of token costs. Any upstream change (model routing, cache blow-up) cascades to consumer pricing.
  • Human divide: Token pricing creates a societal divide between those who can afford advanced AI and those who cannot. Enterprises route certain teams to cheaper models, while startups may receive millions in tokens to disrupt incumbents.
  • No simple caps: One Fortune 100 executive advises against crude caps; instead, talk to outlier users because they might be doing something innovative.

Deep dive: Token economics in practice

Tokens serve multiple roles simultaneously: the unit of output from hardware, the way labs price inputs and outputs, and the value unit enterprises seek to monetize. This abstraction simplifies pricing for hyperscalers but obscures enormous complexity. As SAP's FinOps team notes, "You pay per token, and this little token hides an enormous complexity underneath predictability."

Storment describes three phases of AI usage: the old days before ChatGPT, the good old days when chatbots could write decent code, and the post-November 2025 world of highly capable models. In the good old days, users enjoyed all-you-can-eat tokens and subscription models. Token leaderboards celebrated high usage. Today, no one can afford to waste tokens. Amazon senior VP Dave Treadwell pleaded, "Please don't use AI just for the sake of using AI."

Between June and November 2025, global token usage followed a linear path. Then new models and agentic patterns entered the market. Context windows jumped from thousands to millions of tokens per conversation. Agentic systems added loops, retries, and corrections, causing costs to skyrocket. Companies that subsidized that behavior soon faced massive bills. For example, SemiAnalysis estimated that a $200 Anthropic plan gave $8,000 worth of Claude tokens, while a similar OpenAI offering gave $14,000 worth of Codex tokens. Those days are over.

"So now what matters more than anything is AI value," Storment told the FinOps X audience. "We've got to bring value back to what we're doing… We're in an era where tokens are the main measurement. We're in an era where tokens are in everything in software, and they're driving a lot of the global token economy."

Scarcity and the Jevons paradox

If Moore's law and hyperscale competition were the only forces, token prices would keep falling. Indeed, since 2023, they have dropped dramatically. SAP's internal telemetry shows a clear downward trend in cost per token. However, the floor may be in sight. Storment notes that since November 2025, token prices have been flat, linked directly to hardware and power constraints: "We can't get enough hardware, we can't get enough power… we're seeing backlogs, we're seeing long commitment periods, and we're seeing shortages." Intel's CEO expects no real relief in GPU supply until 2028. SAP's Frederik Pohl warns of supply chain constraints and rising hardware prices for new frontier models.

The net result is a classic Jevons paradox: falling unit cost leads to exploding total spend. At SAP's scale, unit costs fell, but spend doubled in some months. Goldman Sachs projects global usage rising from 6 quadrillion tokens today to 120 quadrillion in 3.5 years. Even if token prices drop further, they are unlikely to fall 24x as fast as volume grows.

How SAP built an AI FinOps framework

When SAP first looked for AI cost data, they hit a wall. Existing cloud tools could show total spend per provider but not per model or per use case. So they manually pulled and merged data from multiple tables. The resulting picture transformed the conversation: within days, the CTO demanded regular updates. Pohl noted, "If you have a CTO asking for a number, that's not a question, it's a mandate."

SAP's framework rests on three pillars:

  • Spend visibility: Track what is consumed, how, and where, across models, platforms, business units, and regions.
  • Economics: Measure efficiency with token-level metrics such as input/output ratios, cached token ratios, and token-to-spend drift (to see if costs rise due to volume or mix shifts to pricier models).
  • Value: Connect spend to business outcomes via cost per use case and inference cost by revenue. This reveals which AI features are economically viable and whether product margins work.

Pohl emphasizes that every token must earn its cost, echoing Nvidia CEO Jensen Huang's phrase "token factory effectiveness." The factory spans silicon, data center leases, model routing, and prompt design.

Tokenomics: from FinOps to full lifecycle

Tokenomics goes beyond cost control to the full lifecycle of tokens as an economic good. Storment defines it as "the emerging discipline of converting energy and capital into AI tokens and resources, consuming those tokens and all the related technology to drive efficient intelligence, and then ultimately drive value on the backend." This breaks into production (energy and capital to create tokens), consumption (allocation, forecasting, optimization), and value (monetization, pricing adjustments, labor implications).

This directly collides with SaaS business models. For example, Microsoft's GitHub Copilot shifted toward explicit usage-based charging, angering developers who relied on unlimited tokens. Labs also tighten screws invisibly: Anthropic's Fable model card would silently drop users to a different model under certain conditions (later walked back). Such policies make naive cost-per-token metrics useless. Storment notes, "A token can cost two cents per million, or it can cost 35 per million, just from a cost perspective," and even at the same rate, one token may drive far more value than another depending on use.

Business models evolve under token pressure

Most customers will never see a line item labeled "120 quadrillion tokens." Instead, vendors build layers: credits and opaque consumption (like putting quarters in a machine), hybrid subscription plus usage (predictable base with token-denominated overages), or direct pass-through models (honest token meters wrapped in dashboards). All are vulnerable to upstream shocks—changes in the token factory, model routing, cache efficiency, or forecasting errors—which can force changes in consumer pricing and cascade through banks and other industries.

The Linux Foundation is spinning up a Tokenomics Foundation alongside the FinOps Foundation to create vendor-neutral specifications for measuring and allocating token costs. The FinOps Focus specification is already being extended for token-level telemetry, and a new program will certify providers' billing pipelines.

The human side of token pricing

Token pricing shapes who gets to use powerful AI. Storment sees a societal divide between those who can afford AI and those who cannot. At the enterprise level, some teams get the latest model, others are routed to cheaper models. Yet crude caps can be counterproductive: one Fortune 100 executive advised against shutting down outliers; instead, talk to them to discover innovative uses.

For individuals, especially new workers, token pricing feeds anxieties about AI and jobs. Storment's view is nuanced but stark: "I don't think AI is immediately coming for everybody's job, but I think the person who's better at AI is coming for the job of the person who's not using AI." If token prices restrict who can learn and experiment, the divide will deepen.

For both companies and individuals, we are moving quickly into an AI-token-based economy. This will lead to a far more expensive AI world. The one thing we know for certain is that costs will be orders of magnitude higher than they have been.


Source: ZDNET News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy