Sitemap

Stop Measuring AI Productivity by Tokens

5 min readJul 1, 2026

--

Press enter or click to view image in full size

More AI usage does not always mean more business value.

Does using more AI make a company more productive?

Not necessarily.

It is easy to measure AI activity. Leaders can see active users, API calls, prompts, tokens, and tool usage almost immediately. Those numbers look clean on a dashboard, which makes them tempting to use as proof that an AI rollout is working.

But token volume is not productivity.

A token is simply a unit of text processed by a model. Prompts, outputs, cached context, and reasoning steps can all contribute to token usage and cost. That makes tokens useful for understanding consumption. It does not make them a reliable measure of value.

This is where “tokenmaxxing” goes wrong.

Tokenmaxxing treats higher AI usage as an achievement in itself. More prompts. More model calls. More tokens consumed. Sometimes even internal leaderboards or incentives built around usage.

The problem is obvious once you say it plainly: a team can burn through a lot of tokens and still produce very little useful work.

As more companies try to connect AI investments to business outcomes, the better conversation is moving from token spend to real P&L impact: what changed in productivity, quality, revenue, cost, risk, or customer experience?

Tokens Are Inputs, Not Outcomes

Tokens are closer to fuel than distance traveled.

They tell you something was consumed. They do not tell you whether the organization got anywhere.

A team might generate long documents that nobody uses. A coding agent might loop through repeated attempts without producing deployable code. An analyst might ask the same question five different ways and still fail to reach a decision.

From a dashboard, all of that can look like adoption.

From a business perspective, it may just be waste.

This does not mean token data is useless. It matters for cost management, capacity planning, model selection, and understanding where AI is being used. But it should support ROI analysis, not replace it.

The better question is not, “How much AI did we use?”

It is, “What valuable outcome did that AI usage produce?”

That is why some teams are starting to think less about raw usage and more about tokens per approved outcome: how much AI consumption is required to produce work the business actually accepts?

AI Productivity Is Real, but It Is Uneven

Rejecting tokenmaxxing does not mean dismissing AI productivity.

There is strong evidence that generative AI can improve performance in real workflows. A major field study found that generative AI increased issues resolved per hour by 14% for customer support agents, with especially large gains for newer and lower-skilled workers. The study was first circulated by NBER and later published in the Quarterly Journal of Economics as Generative AI at Work.

That detail matters.

AI did not affect every worker the same way. It helped some groups more than others. Stanford HAI summarized the same finding clearly: AI productivity gains can be largest for workers who are not already top performers in the task.

That is why companies need to measure AI at the workflow level, not only at the company-wide usage level.

Good AI measurement starts with a baseline.

How long did the task take before AI?
How many cases were resolved?
How much rework happened?
What was the error rate?
What did customers experience?
Did quality improve, stay flat, or decline?

Without those answers, a company may only be measuring enthusiasm.

A Better AI ROI Scorecard

A stronger approach is to measure AI across five levels.

1. Adoption
Who is using AI? Which teams have active use cases? How often is it being used? Adoption is useful, but it is only a leading indicator.

2. Efficiency
Did the workflow get faster? Did cycle time drop? Are support cases, reports, tickets, or reviews moving through the system more quickly?

3. Quality
Did accuracy improve? Did defects, rework, escalations, or compliance issues decrease? Did customer satisfaction hold up?

4. Business outcomes
Did AI help generate revenue, avoid costs, retain customers, accelerate deals, reduce risk, or improve service?

5. Unit economics
What does it cost to produce one accepted output, resolved case, completed task, or successful deployment?

This structure keeps teams from jumping from “people used the tool” to “the investment worked.”

McKinsey makes a similar point in its work on how companies can measure and realize the full value of AI: value measurement needs to connect adoption and operational impact to financial outcomes. Usage alone is not enough.

Adoption is only the beginning. ROI comes from what happens after adoption.

The Simple ROI Formula

A basic AI ROI formula still helps:

AI ROI = (Financial Benefits − Total AI Costs) ÷ Total AI Costs × 100

The hard part is defining both sides honestly.

Financial benefits may include recovered labor capacity, faster delivery, fewer errors, reduced outsourcing, additional revenue, or avoided hiring.

Total costs should include more than tokens. Companies also need to account for subscriptions, infrastructure, implementation, training, governance, human review, maintenance, security, and rework.

There is also one important caveat: time saved is not automatically money saved.

If AI saves a team ten hours, but those hours do not turn into more customers served, faster releases, better decisions, or lower costs, the business value is still theoretical.

Recovered capacity only becomes ROI when the organization uses it well.

From Tokenmaxxing to Value Maximization

The goal is not to use the most AI.

The goal is to get the most value from the AI you use.

That means matching each task with the least expensive model that can do it reliably. It means caching reusable context, limiting unnecessary agent loops, comparing frontier models with smaller alternatives, and setting quality thresholds before deployment.

It also means budgeting around use cases, not abstract consumption.

A support assistant should be measured by resolved cases, escalation rates, handle time, and customer satisfaction.
A coding assistant should be measured by accepted pull requests, review time, defect rates, and cycle time.
A research assistant should be measured by accuracy, source quality, time saved, and decision usefulness.

Tokens still matter.

They just should not be the scoreboard.

Measure the Distance, Not the Fuel

Token consumption is like fuel burned by a vehicle.

It helps leaders understand operating cost. It does not show whether the vehicle took the right route, reached the destination, or delivered anything valuable when it arrived.

The next phase of enterprise AI needs a better measurement culture.

Instead of asking, “How much AI are employees using?” leaders should ask:

What work improved?
What quality changed?
What cost was avoided?
What revenue was influenced?
What risk was reduced?
What did the customer experience?

That is how AI moves from activity to value.

Organizations that make this shift will have a clearer view of what is working, what is wasteful, and where AI deserves more investment.

The real goal is not more tokens.

It is better outcomes.

--

--

ODSC - Open Data Science
ODSC - Open Data Science

Written by ODSC - Open Data Science

Our passion is bringing thousands of the best and brightest data scientists together under one roof for an incredible learning and networking experience.