The same companies that ran token-use leaderboards just months ago are now capping how much AI their employees can touch.
Tech firms have flipped from pushing maximum AI use to actively limiting it, because the bills got brutal. Meta, Uber, Walmart, and Amazon are all reining in spending after costs exploded, with Uber admitting it burned through a full year of projected AI budget in just four months. The new mantra: use less, measure better.
Picture an engineer earlier this year, racing up an internal leaderboard for burning the most AI “tokens,” cheered on by a company that told everyone to use as much artificial intelligence as humanly possible. Now picture that same leaderboard quietly deleted, a monthly usage cap in its place, and a finance team asking what all those tokens actually shipped. That swing happened in a matter of months.
What Happened
Earlier this year, the message from tech companies was blunt: use as much AI in your work as possible. Employees called it “tokenmaxxing,” a token being a unit of AI use roughly equal to a word fragment. Staff at Meta and Amazon even competed on leaderboards tracking who burned the most.
Then the invoices arrived from AI providers like Anthropic and OpenAI, and they were not cheap. Meta told employees last week it would soon limit AI use after an “exponential increase” in costs. Uber said in May it had blown through its projected annual AI spending in just four months and placed monthly caps on AI coding tools. Walmart set its own limits. Amazon and Meta pulled down the leaderboards entirely.
The new buzzword, per Rob May, CEO of AI consultancy Neurometric and author of “The Tokenminning Manifesto,” is “tokenminning,” short for token minimizing. Use is in flux, and nobody has fully figured out the rules yet.
The Backstory
Here is why the meter runs so fast. OpenAI and Anthropic sell subscriptions from $10 to $200 a month, but the real money comes from enterprises like Meta, Shopify, and Amazon, which pay both subscription fees and for every token their tens of thousands of workers consume. More tokens used, more money owed.
The trouble is that tasks got heavier. Summarizing a meeting transcript might cost a few hundred tokens. Writing code for a new feature can cost tens of thousands. Engineers also stopped just chatting with bots and started deploying AI “agents” that grind on complex tasks for hours, which means a single engineer can torch tens of thousands of dollars in tokens a month. And many defaulted to the most powerful model for everything, even though cheaper ones exist. Anthropic’s newest model, Fable, costs twice as much as its predecessor, Opus.
The Plan
The fix companies are converging on is simple: reserve frontier models for the hard stuff, and swap in cheaper models everywhere else. AT&T’s chief AI officer, Andy Markus, says firms can save as much as 90 percent by opting for less advanced models, noting that “for most use cases, the latest greatest frontier model isn’t needed.”
The deeper shift is in how success gets counted. Salesforce CEO Marc Benioff says his company still plans to spend hundreds of millions on AI this year, but now tracks “agentic work units” instead of tokens, a metric meant to measure output rather than raw usage. Meta, on track to spend billions, says it wants to “find places we can spend less while getting similar or better business results,” and has told engineers to use its internal MetaCode assistant over third-party tools where possible.
The Business Model Angle
This is a clean case of Goodhart’s Law: when a measure becomes a target, it stops being a good measure. CEOs who could not gauge AI skill reached for the easiest proxy, token volume, and got exactly what they incentivized, volume over efficiency. The leaderboard didn’t reward smart AI use. It rewarded burning tokens.
The lesson for operators runs deeper than AI. Any time you pay per unit of consumption but reward your team for consumption itself, you have built a machine that prints cost with no guaranteed link to value. The companies winning the next phase are the ones replacing a usage metric with an output metric, as we unpacked in our earlier breakdown of the corporate scramble to control AI token bills. If you cannot draw a straight line from spend to shipped value, you are not measuring performance. You are just measuring appetite.
The Risk
The honest counterpoint: nobody is sure tokenminning actually works yet, and cutting too hard could cost more than it saves. Uber COO Andrew Macdonald put the tension plainly, saying that if you cannot connect spending to useful features shipped, “that trade becomes harder to justify. That link is not there yet.” The problem is that the link being weak today does not mean the spending is wrong, it may just mean the measurement tools are immature.
There is also a real chance companies overcorrect, throttling the experimentation that produces breakthroughs in the first place. And the effect on Anthropic and OpenAI is genuinely unclear. At the height of tokenmaxxing, both reported record revenue from coding tools. If enterprises get disciplined and shift to in-house tools like MetaCode, that growth engine could cool. Cheaper, smarter usage is good business hygiene, but swung too far it becomes the corporate version of skipping lunch to save money, then crashing at 3 p.m.
Quick Questions
What is tokenmaxxing, and why did it end?
It was the early-2026 push to use as much AI as possible, with some companies running internal leaderboards for token use. It ended when the bills landed and firms realized they were paying for volume, not results.
How much can companies actually save?
AT&T’s chief AI officer says up to 90 percent on many tasks, simply by using cheaper models instead of defaulting to the most powerful one for everything.
Does this mean Big Tech is spending less on AI overall?
No. Meta is still on track to spend billions and Salesforce hundreds of millions this year. They want the same or better results for fewer tokens, not a smaller AI bet.
What is replacing token count as the metric?
Output-based measures. Salesforce now tracks “agentic work units” meant to reflect work done, not tokens consumed, signaling a shift from measuring usage to measuring value.
The Business Model Analyst Take
The story here is not “AI got expensive.” It is that the entire industry picked a vanity metric, optimized hard for it, and is now paying to unwind the incentive. For founders and operators, the takeaway is sharp: pick your success metric before you scale the behavior, because whatever you reward, you will get a flood of. Tie spend to shipped value from day one, keep a cheap default and reserve the expensive tools for problems that truly need them, and treat any usage-based cost as a variable you actively manage, not a subscription you forget about. The companies that thrive in the agentic era won’t be the ones that used the most AI. They’ll be the ones that knew what every token bought.
