Tokenmaxxing Is a Vanity Metric. Don't Let It Kill Your AI Strategy.

An efficient AI stack beats an impressively large one every time.

Written by Eugene Cheah
Published on Aug. 25, 2026
A hand puts a token into a subway turnstile coin slot
Image: Shutterstock / Built In
Brand Studio Logo
REVIEWED BY
Summary: The tech industry’s obsession with tokenmaxxing and giant context windows creates a false measure of productivity. Counting raw tokens reflects only computational cost, not actual business value. To maximize ROI, companies must shift from monolithic models to specialized, smaller AI models.

The discussions taking place within the tech community lately, and especially after the latest Google I/O event, make it quite clear that we love our scale. Big companies and developers are touting the benefits of incredible throughputs, enormous data output rates and ridiculously long context windows as part of some tokenmaxxing process. If you haven’t heard, tokenmaxxing is the process of intentionally maximizing the consumption of AI tokens to show higher usage. In turn, the hope is this consumption leads directly to proportionally greater productivity. And from an aesthetic point of view, it all sounds rather exciting when presented in a presentation slide. 

But all of this shows just how little we understand about the process through which enterprise software generates real business value. Counting raw token numbers is an awkward indicator of performance. The tokens only show the computational cost that the model used, not the correctness or usefulness of the output.

As LLMs have gone from chatbot to agents, how we measure their progress and the work they do has also changed. These days, the most effective way to contextualize token pricing is cost per successful task. Instead of thinking per conversation and per tool call, think bigger. Realize that agents that can go out and solve tasks over a long period of time should be measured accordingly. The 2026 BCG AI Radar Survey has revealed that organizations plan to allocate an average of 1.7 percent of their total corporate revenue to AI initiatives this year. Yet that same data reveals a staggering paradox: 60 percent of those executives report seeing minimal or no material financial return from these investments. Relying on excessive computational effort as a primary means for doing business does little to reflect actual business value.

What Is Tokenmaxxing? Why Is It Bad?

Tokenmaxxing is the practice of intentionally maximizing AI token consumption to demonstrate higher usage, with the goal of showing greater productivity. Here’s why you should avoid it:

  • Measures Cost, Not Value: High token counts measure computational effort and fuel used, not the accuracy, usefulness or quality of the output.
  • Creates Vanity Metrics: Large context windows and oversized outputs act as vanity metrics that often lead to low return on investment (ROI).
  • Increases Costs and Latency: Generating excessive, unnecessary tokens inflates computing expenses, increases processing latency and degrades the overall user experience.

More on AI + FinanceThe Era of Affordable AI Is Over. What Comes Next?

 

Not All Tokens Are Created Equal

Software consumers have been bombarded by endless claims of unlimited capacity, but very few of them seem to wonder what the pieces of information that we’re generating accomplish. The truth of the matter is that not all tokens are created equal. Computational tasks deliver wildly varying degrees of business value depending on the nature of the action performed. But only the tokens that go into conducting an exact query or solving a crucial bottleneck in a customer’s problem have high business value.

Conversely, chasing massive, unoptimized outputs is an exercise in diminishing returns. Recent research from Forbes underlines this gap, revealing that fewer than 1 percent of enterprise executives report a significant AI return on investments of 20 percent or more. Instead, the vast majority find themselves trapped with marginal gains of just 1 to 5 percent.

High throughput and oversized context windows are typically vanity metrics that technical teams use to prove their AI models capabilities without any practical implications. In reality, for a company operating in the world of finite budgets, having to provide input for a whole book’s worth of data just to produce one small summary paragraph would be an absolute waste of resources rather than an engineering marvel. Return on investment always comes from efficiency.

 

Data Production Is a Cost Center

We need to start treating AI data production as the cost that it is. Each unnecessary token produced by AI becomes another burden for the processing systems, leading to latency increases and harming user experience. In terms of economics, such practices simply eat away at margins, and computing is a costly process that you can’t ignore.

Therefore, if a company decides to incorporate any enterprise workflow based on producing useless data, it would fail to scale efficiently. In one such case, an internal misconfiguration at Amazon led to a disastrous cloud cost of $1.8 million because an agent was performing long, repetitive data-matching tasks without budget limits. Related to this is a claim by industry analysts that companies are shelving AI programs entirely due to the high cost of moving unstructured, unoptimized data.  

We got here thanks to lazy architectural principles, which involved using one massive and monolithic model for all of a company’s needs. Being overly generalist by nature, this approach leads to the inefficient overgeneration of text. Running a large, pre-trained model to perform a simple task like classification or straightforward information extraction amounts to driving a semi truck to the nearest supermarket to purchase a single carton of milk. Technically, it works, but it would take up too much gas. These models require a substantial number of tokens and generate lots of text to make their final prediction.

More on TokenomicsWhat Is Tokenomics, and Why Is Everybody Talking About It?

 

Say No to Tokenmaxxing

The future will see a different trend in AI usage. Instead of focusing on using the biggest models in the market, we should be shifting our attention towards minimizing token usage. Modern approaches tend to use highly specialized, smaller models that would perform more efficiently within a particular application in a business. If you train a nimble model to work within a certain function in your company, whether legal services, programming assistance or call center management, you’ll benefit greatly from increased accuracy and reduced compute power.

For example, if you applied a large frontier model on simple tasks such as identifying invoice numbers from a PDF or routing help desk tickets, you would waste a large amount of compute resources. It’s best to save the large models for complex and unstructured tasks such as synthesizing multiple financial reports. A good rule of thumb for allocating compute is to look at volume and complexity of each task that needs to be completed. If a task is high-volume and straightforward, use a small and specialist model. Save the expensive large models for low-volume and complex tasks that carry higher business value. 

It can be tempting to ignore such inefficient operations during the excitement of the launch of a new product. When a new, high-profile AI tool comes out, it’s hard to avoid generating a huge number of tokens by using it. At that time, prices will still be low with many cloud credits available, so the usage cost will remain relatively low, hidden behind this discount-heavy period. Only once the discounts are removed and real costs are clear can you assess the sustainability of that project’s margins.

To avoid falling into this trap, build in guardrails from day one. Try to forecast your costs using full, undiscounted prices and set hard limits like automated budget alerts and token caps. Most importantly, as explained above, be smart about which models you use where to make sure you’re paying for actual business outcomes, not just raw compute. 

The time has come for the industry to stop focusing so much on outputs. I expect that the best AI technologies in the near future will be distinguished not by their ability to produce but rather by the minimum amount they will need to deliver results.

Explore Job Matches.