Hyperscalers are accelerating efforts to scale artificial intelligence amid the next wave of “agentic” and “physical” AI applications, and investor focus is increasingly shifting from raw token throughput to token economics. Pershing Square Capital Management CEO Bill Ackman laid out a thesis in which lower per-token costs could expand demand for AI models, lift cloud usage, and improve profit margins—assuming technology improves efficiency faster than markets expect.
The core argument centers on whether the industry can generate more AI tokens per unit of electricity, potentially reducing costs tied to data-center processing and Nvidia-powered infrastructure. Ackman’s view also points to why major cloud providers—especially Microsoft and Alphabet—have emphasized high-volume token processing as they plan for sustained, capital-intensive investment cycles.
Key takeaways
- Price move: No specific market move was described in the provided material.
- Catalyst: Bill Ackman’s thesis that improvements in token-cost efficiency could justify continued heavy hyperscaler capital spending.
- Key implication: If token costs fall, hyperscalers may serve more AI demand at better margins, supporting revenue growth tied to cloud and model usage.
- Investor lens: Markets may be underestimating the profit opportunity from optimizing cost per generated token.
Why token costs are becoming a central AI metric
According to the material, AI systems such as ChatGPT process user requests as “tokens,” and those tokens are not limited to text prompts. Tokens are also produced when users generate images, videos, and other types of content. In practice, each token requires electricity to compute, which makes token economics tightly linked to data-center efficiency and the semiconductor infrastructure that powers AI workloads.
The report described a supply-chain relationship in which cloud providers purchase Nvidia chips to generate tokens at scale. It also noted that other AI chipmakers participate in token generation, but that Nvidia holds the largest market share, making token cost improvements particularly important for how hyperscalers manage compute spending.
Investors, the piece said, watch how many tokens can be generated per watt of electricity consumed as a proxy for future cost declines. It outlined a simple efficiency framework: if tokens per watt were to double, token-related costs could fall proportionally. Ackman’s position, as presented, is that hyperscalers are likely to optimize this metric over time, and that the market has not fully priced in how those improvements could affect profitability.
How lower token costs could change hyperscaler economics
The material argued that token optimization would allow hyperscalers to deliver the same level of customer service with fewer chips, supporting higher operating margins. However, it also emphasized that the impact may extend beyond cost savings to demand growth—an important distinction for investors evaluating whether AI spending will translate into durable earnings power.
According to the discussion, lower token costs could increase demand for AI models and expand consumption of cloud platforms. The piece cited the Jevons Paradox, which holds that improved efficiency can spur greater usage. In this context, the concern that falling token costs might compress AI compute demand was reframed: even with improved efficiency, overall demand for AI chips could rise if users and workloads scale in response to more affordable compute.
The report also highlighted that token optimization matters as agentic AI and physical AI expand, because these applications can increase how extensively systems interact with users and environments. If hyperscalers can run those workloads more efficiently, the expected payoff from ongoing infrastructure investment could become more visible in financial results.
What companies’ token-volume disclosures suggest
In its thesis, the material pointed to internal usage figures shared by major technology firms. It said Microsoft processed more than 100 trillion tokens in a single quarter in 2025, describing that as a fivefold year-over-year improvement. It further claimed that Microsoft processed 50 trillion tokens in one month within that period, indicating meaningful month-over-month growth.
For Alphabet, the report stated that the company told investors it was processing more than 16 billion tokens per minute in Q1. The takeaway offered was that large-scale processing continues to grow, which can help validate that hyperscalers are not just investing in capacity but also finding demand strong enough to keep utilization climbing.
Bigger picture: $700 billion in capex and the path to returns
According to the material, hyperscalers are approaching $700 billion in AI spending in 2026. For investors, the key question is not only whether that capital expenditure cycle will expand capacity, but whether improved token efficiency can convert that spending into higher returns.
As presented by Ackman’s thesis, token-cost improvements could support both sides of the earnings equation: reduced compute cost per unit of AI work and increased usage that expands the addressable market for AI models and cloud services. The report’s conclusion is that hyperscaler optimization efforts could eventually remove some of the uncertainty around when heavy AI investments translate into higher profits and faster revenue growth rates.
Investors may want to watch for future disclosures and guidance that connect AI workload growth with efficiency metrics—especially data center performance, cost trends tied to inference and training, and updates on how agentic and physical AI products are driving usage. With hyperscalers continuing large-scale investment, the next evidence point is likely to come through quarterly reporting, investor presentations on utilization and costs, and broader signals from the semiconductor supply chain and data-center buildouts.







