The AI Bubble Debate Is Asking the Wrong Question. The Real Story Is Margin Migration

AI Bubble Debate

Michael Burry is shorting Nvidia and Palantir. Ray Dalio says his bubble indicators are flashing at 1929 and 2000 levels. Berkshire is sitting on a record cash pile while four companies commit to spend roughly $700 billion on AI infrastructure this year against an estimated $12 billion in consumer AI revenue. The “is it a bubble” debate has swallowed the financial press whole.

It is also the least useful question you can ask right now.

The fight over whether AI stocks are in a bubble is really a fight over a single binary: pop or no pop. But the thing actually reshaping the AI economy in mid-2026 is not a demand cliff. Demand is still compounding. What is changing is who captures the margin on that demand. The model layer is commoditizing, corporate buyers have flipped from “use as much AI as possible” to “justify every token,” and the profit pool is migrating down the stack from frontier model providers toward compute, cloud, and orchestration. That is a business-model story, not a bubble story, and it has clearer near-term consequences than any 1929 analogy.

What the bull-bear fight is actually about

Strip away the theatrics and the bears and bulls barely disagree on the facts. Both sides accept that AI is real and transformative, that hyperscaler spending is unprecedented, and that the timing of when revenue catches up to capex is genuinely uncertain.

The bears focus on that last gap. Burry has compared the 2026 trade to the final months of the dot-com peak and warned that “circular funding” among the big players could later look like a picture of fraud rather than a flywheel. Dalio, on Bloomberg in early June, put it more gently: “Great technologies create bubbles.” Ed Zitron and Jeremy Grantham round out the chorus. Their shared claim is narrow: not that AI is fake, but that current prices do not reflect the risk that revenue arrives later and smaller than the buildout assumes.

The bulls, including Satya Nadella, Sundar Pichai, and Bill Ackman, point at the cash already moving. Nvidia just reported roughly $81.6 billion in quarterly revenue, up about 85% year over year, with its data center segment near $75.2 billion. Microsoft’s cloud RPO jumped 99% to $627 billion. These are not the financials of a hollow mania.

Here is what both camps miss by arguing about the index level: the most important repricing of 2026 is happening inside the AI stack, not to it.

Demand is not the problem

Start with the part the bears get wrong on timing. AI demand is not softening. It is doing the opposite, and it is doing it through a mechanism that punishes anyone betting on a usage collapse.

When token prices fall, usage rises faster than the price drop, which is Jevons paradox in its purest form. The clearest proof point is the one everyone got backwards. In January 2025, when DeepSeek showed a Chinese lab could match Western reasoning cheaply, investors concluded that cheaper intelligence meant collapsing chip demand and wiped roughly $600 billion off Nvidia in a single session. The thesis was exactly wrong. Cheaper intelligence expanded compute demand. Nvidia is now worth more than $5 trillion, up about 50% over the year.

The forward numbers say the same thing. Goldman Sachs projects enterprise token consumption rising 24 times by 2030, to around 120 quadrillion tokens a month. Hyperscaler backlogs, the contracted-but-not-yet-delivered demand that shows up as Remaining Performance Obligations, keep climbing: Google Cloud’s backlog jumped 93% in a single quarter to $462 billion. The demand side of the equation is not where the cracks are.

The tokenmaxxing reversal

The cracks are in how that demand gets paid for. Earlier in 2026, the corporate posture toward AI was “tokenmaxxing”: use as much as humanly possible, run internal leaderboards for token consumption, treat the spend like brand marketing where the upside is unmeasurable but obviously good. The market rewarded the attitude.

That ended fast. As we covered when the shift broke, a string of the biggest names flipped from maximizing AI use to actively rationing it once the bills landed. The trigger was the move by Anthropic and OpenAI to shift customers from flat subscriptions to token-based billing, which exposed buyers directly to the cost of every prompt and every autonomous agent. The numbers got brutal:

CompanyThe moveWhat broke
Uber$1,500/month per-tool capBurned its entire 2026 AI budget by April
WalmartToken caps on internal “Code Puppy” agentAdoption surged past budget
AmazonShut down its internal usage leaderboardEngineers ran junk prompts to climb it
MetaToken governance, mid-JuneCosts approaching billions internally
Cisco, AT&T, MicrosoftUsage restrictionsPer-employee costs into the thousands monthly

The quote that captures the turn came from Uber COO Andrew Macdonald: the link between token spending and shipped product value “is not there yet.” A KPMG survey found only 26% of companies have comprehensive visibility into their AI costs. That opacity is now closing, and closing visibility means buyers can finally measure output against input, which means spend gets disciplined.

This is the shift from brand-marketing logic to performance-marketing logic, applied to AI. And performance measurement is never kind to the most expensive supplier.

The model layer is commoditizing

Which brings us to the supplier problem. The frontier model, once the moat, is becoming the commodity input.

The AI Bubble Debate Is Asking the Wrong Question. The Real Story Is Margin Migration

The gap between open-weight and closed frontier models has compressed from double digits to low single digits on the benchmarks that track real work, and it compresses further with each quarterly release. DeepSeek V4 Pro matches GPT-5.5 and Claude Opus 4.8 on most agentic benchmarks at roughly 10 to 13 times lower cost per output token. Zhipu open-sourced GLM-5.2 on June 13 under an MIT license with a million-token context window. By May 2026, Chinese open-weight models accounted for roughly 61% of all tokens consumed on OpenRouter, the largest neutral model router, and four of the five most-used models were Chinese. Meta’s Llama, the open-weight leader two years ago, fell off the rankings entirely.

That cost gap is the engine of everything downstream.

The strategic point is not “China is winning AI.” It is narrower and more useful: the open model layer has been commoditized, and that repricing looks permanent. Weights are becoming a customer-acquisition funnel for cloud platforms, not a standalone business. When most enterprise work (document analysis, code review, classification, extraction) runs fine on a model at a tenth to a thirtieth of frontier cost, the rational default for that work is no longer the frontier.

Where the margin actually goes

So if intelligence is getting too cheap to meter, where does the value accrue? Down the stack and out to the edges.

Inference, the act of running trained models, now represents roughly two-thirds of all AI compute, up from one-third in 2023. Even a free open-source model is too large to run on a laptop, so it has to be hosted somewhere, which means a cloud provider collects margin on every token regardless of which model wins. That is why infrastructure capital is flooding into the serving layer: Baseten raised $1.5 billion at a $13 billion valuation, and OpenRouter, which routes inference across providers, is moving roughly 25 trillion tokens a week.

The pattern rhymes with utility economics. The companies that built the railroads and the early power grids did not capture the value of everything that ran on them. They earned the regulated margin on transmission. In AI terms:

LayerPosition in 2026
Frontier LLM providersMoat survives only at the bleeding edge; pricing power eroding everywhere else
Cloud / inference hostingCollects margin on every token, open or closed, while compute stays scarce
Chips and siliconDemand expands as cheaper models pull more usage (Jevons)
Orchestration / agentsWorkflows and proprietary data graphs become the defensible asset

The frontier still has a real moat, but it has shrunk to the literal bleeding edge of novel capability. As Nadella himself argued in a June essay on “token capital,” firms that only rent models risk having their expertise commoditized, while firms that own a learning loop compound an advantage. The model is the input. The advantage sits above and below it.

Why this matters more than the bubble call

Markets do not trade the first derivative. They trade the second.

What moves a stock is usually not whether growth is positive but whether the pace of growth is accelerating or slowing. After a long run where the momentum factor crushed everything else, a slowdown in the second derivative tends to drive reallocation toward quality and away from the most stretched narratives. That is the real near-term risk in the AI complex, and it has nothing to do with whether the word “bubble” applies.

Here is the chain of logic that actually matters for the next few quarters. Corporate budgets finding a ceiling means more optimization. More viable open models means more competition and thinner margins for LLM providers. Thinner model margins cap what providers will pay for compute, which caps what chipmakers can charge up the supply chain, which dictates how much capex providers are willing to commit. Margin drove the entire AI complex up. Margin is also what can drive it back down. The ROI has to make sense.

The wrinkle is that commoditization is not uniformly bad. It is bad for frontier model providers and good for cloud providers, at least while compute stays constrained, because the cheaper models still need somewhere to run and the hosting margin survives the price war. The “more competition equals more usage equals more compute demand” loop is real. The phrase doing all the work, though, is “all else equal,” and it never is. Datacenter supply coming online, memory pricing, and demand destruction from rising hardware costs all cut the other way.

What this means if you run a business, not a portfolio

For operators rather than traders, the read-through is concrete and worth acting on now, well before the headlines catch up.

First, your AI cost structure is about to become a competitive variable, not an afterthought. The companies that capped spending did not do it because AI stopped working. They did it because they could not measure it. Build the measurement before you build the dependency.

Second, route by task. Not every workload needs the frontier model, and the gap between “good enough” and “best” has collapsed for the majority of routine work. Reserving frontier models for the genuinely hard 20% and pushing the rest to cheaper options is now how disciplined teams keep token costs roughly flat while usage grows.

Third, if your product’s only value is wrapping a frontier API and marking it up, the margin underneath you is evaporating. The defensible position is the workflow, the integration, and the proprietary data that no model can replicate, not the intelligence you rent.

The broader AI buildout, meanwhile, keeps showing up in places that have nothing to do with your software bill. As we have written, the scramble for the same chips, cables, and power is quietly lifting prices on everything from consoles to electricity. The infrastructure story and the margin story are two halves of the same shift.

The Business Model Analyst Take

“Is AI a bubble” is a question designed to be unanswerable, which is exactly why it generates so many headlines and so little insight. Nobody can call the top, the bears have been early before, and the market can stay irrational longer than any short can stay solvent.

The more honest and more useful framing is this: the AI economy is repricing internally, and the repricing has a direction. Demand is compounding, but the model layer that everyone built valuations around is commoditizing, and the margin is migrating toward whoever controls scarce compute and the workflows on top. That can be true whether or not stocks correct. It tells you where the durable economics are settling regardless of what the index does next week.

The bubble debate asks whether the music stops. The better question is who owns the room when it does. Right now, the answer is moving from the people who make the models to the people who host them and the people who build real workflows on top. Position accordingly, and stop waiting for a pop that may or may not come.

This is not financial advice. It is a read on where business-model value is migrating in the AI stack.

UNLOCK THIS FREE DOWNLOAD

DOWNLOAD NOW

Fill Your E-mail to Receive this Download Directly in Your Inbox.

RECEIVE OUR UPDATES

The Biz Model Club

Get daily, no-fluff insights on the latest business models, startup strategies, and trends delivered straight to your inbox.