Everyone has already buried it. But before we shovel the last of the dirt, I want to check that we’re burying the right body.
Let’s start with the founding assumption of early 2026, and let’s state it as plainly as the people acting on it actually believed it.
Tokens are finitely infinite. There’s technically a limited number of them, in roughly the way there’s technically a limited number of grains of sand, which is to say the limit isn’t your problem.
They cost nothing that shows up anywhere you have to look, and they draw on no resource you can feel.
So, if you accept the premise, the only rational thing to do is to consume as many of them as humanly, and inhumanly, possible.

The two numbers everyone was watching in early 2026
Nobody wrote the premise down quite like that, of course. But it’s usually more honest to watch what people do than to listen to what they say, and what people did was build leaderboards.
A Meta employee set one up on the company intranet and it tracked token consumption across more than 85,000 colleagues. Amazon and others followed. The theory behind it was simple enough to fit on a sticky note. More tokens burned meant more AI-powered work getting done.
So whoever sat at the top of the leaderboard was your most productive engineer, and a company climbing the aggregate was a company pulling ahead.
Consumption was the scoreboard, and people played to win.
The obituary everyone already wroteÂ
Then the invoices arrived. Uber burned through its entire 2026 token budget in the first four months of the year, which its own COO later described as a head-exploding moment. Meta quietly took the leaderboard down. Amazon put out internal guidance that amounted to please stop using AI just to use AI, after reporting that some employees had been spinning up agents to do meaningless tasks purely to keep their numbers up. Microsoft started cancelling Claude Code subscriptions across product divisions. Somewhere in the middle of all this, one company reportedly ran through hundreds of millions of dollars of tokens in a single month.
So, the verdict is in, and it’s unanimous. Tokenmaxxing is dead. It was a vanity metric, it bred dysfunction, it lit money on fire, and the grown-ups have arrived to shut it down. I agree with every fact in that paragraph. What I don’t agree with is the conclusion most people are drawing from those facts, because I think it’s the most convenient one.
And yet the instinct wasn’t stupidÂ
If you genuinely don’t know where a new technology creates value inside your organization, you’re not going to find out from a careful pilot with twelve users. Value tends to show up in the tails, in the workflow nobody predicted, in the one team that quietly reinvented how it works. And to see the tails, you need volume. You almost have to flood the system before you can learn very much about the system.
Mathematicians have a clean name for this. It’s the exploration-exploitation tradeoff, and it’s the thing every multi-armed bandit algorithm is built around.
Picture a row of slot machines with unknown payouts. Early on, the correct move is to explore, to keep pulling arms you know nothing about even when it feels wasteful, because the information you gather now compounds over every decision you make later. Over-spending at the start isn’t a bug in that world. It’s the only way you buy the map to success. So, the 2026 instinct to just use the thing, a lot, right away, wasn’t the mistake. On math, it was arguably the right call.
But a bandit algorithm only works because every pull gets scored. You pull the arm, you watch the payout, you update your belief, and then you pull again.
Explore, measure, learn, repeat. What tokenmaxxing did was pull millions of arms and score none of them.
What actually brokeÂ
Two things broke in sequence, and almost everyone is only talking about the first one.
The first is Goodhart’s Law, the old idea that when a measure becomes a target, it stops being a good measure.
The moment token count went from something you could quietly observe to something you were rewarded for, people did the entirely rational thing and started optimizing the number instead of the work. The engineers running pointless agents weren’t being stupid. They were being obedient to the scoreboard someone handed them. And notice the second-order damage here, because it’s the part that really stings.
Their gamed tokens didn’t just fail to signal productivity. They polluted the pool, so that even the honest tokens got harder to read.

Goodhart’s Law as a loop, where the metric corrupts itself
The second break is quieter, and I think it’s the more important of the two, because it would have bitten us even in a world with no gaming at all.
Suppose every token burned in 2026 had been burned in perfect good faith. You’d still be sitting on an ocean of usage exhaust with no instrument for turning it into an answer. Which prompts actually created value, and which were just expensive noise? Where did all that volume genuinely pay off? Nobody could really say, because nobody had built the layer that watches the pulls and scores them.
We just ended up buying the map-making data and then never drew the map.
We have absolutely seen this movieÂ
Rewind to the early 2010s. The phrase back then was big data, and the promise was that if you simply captured enough of everything, the insight would somehow fall out of the pile on its own. So everyone captured everything. Data lakes filled up. And a great many of them slowly turned into data swamps, because hoarding was never the same activity as understanding, and the pipeline from raw signal to real decision was the hard, unglamorous part that far fewer teams actually bothered to build.
It’s the same shape, just on a faster clock. In both eras the industry mistook accumulation for comprehension. It assumed that volume would metabolize itself into knowledge without anyone building the digestive system. Big data took the better part of a decade to learn that lesson. Tokenmaxxing took about four months, which I suppose we can file under progress.

Tokens burned kept climbing while insight stayed flat
What the survivors are quietly buildingÂ
Look at what the companies that hit the wall are doing now.
- Per-task cost tracking.Â
- Model routing, so the expensive brain only gets called when the cheap one won’t do. Â
- Context discipline, so an agent loads what it needs instead of resending everything on every turn. Â
- Attribution, so you can finally point at a spend and a result in the same sentence.Â
Read that list again, because none of it says use fewer tokens. All of it says know what your tokens are doing. These companies aren’t repudiating the volume. They’re finally bolting on the nervous system that the volume needed all along, which is the scoring layer that turns a million blind pulls into an actual map. They just won’t call it tokenmaxxing anymore, because the word smells like the invoice now.
So no, I don’t think the instinct was wrong. Flooding the system to learn fast was, on the math, a defensible bet. Turning that flood into a scoreboard was the unforced error, and building no instrument to read the flood was the mistake that quietly cost the most.
Ultimately, just like the era of big data, the volume was never really the problem. The blindness was.
Which is the whole reason a control plane matters, and it’s what we’re building toward at Lyzr. It gives every token an outcome to answer for, so volume turns into learning instead of noise.
Book A Demo: Click Here
Join our Slack: Click Here
Link to our GitHub: Click Here

