What We’re Really Measuring

Why Amazon's AI usage targets are producing the gaming behavior they were meant to prevent

Share
What We’re Really Measuring

The Financial Times reported this week that Amazon employees have started “tokenmaxxing,” running their internal MeshClaw agent through unnecessary tasks to inflate their AI usage statistics. The pattern, also documented at Meta, follows a target Amazon set for 80% of their developers to use AI weekly, paired with internal leaderboards tracking token consumption.

Why Amazon's AI usage targets are producing the gaming behavior they were meant to prevent
Envato Elements

Amazon has told employees the data won’t appear in performance reviews. Several employees told the FT they don’t believe that for a second.

We understand the pressure that produced this measurement system. Amazon expects to spend roughly $200 billion in capital expenditure this year, with most of that going to AI infrastructure. When an organization commits that scale of capital, leadership reasonably wants evidence of adoption. Some form of tracking has to exist. Saying “we believe in AI” without measuring use is the kind of espoused commitment that produces nothing.

The impulse is right. The execution reveals a problem we’ve watched repeat across industries for the last two years.

Token consumption is an activity measure dressed up as a value measure. An engineer who solves a difficult problem with a single well-crafted prompt scores poorly. An engineer who spawns ten agents to triage trivial emails scores well. The leaderboard, working exactly as designed, rewards the behavior leadership would say it doesn’t want.

This is the structural problem Chris Argyris named in 1974 as the gap between espoused theory and theory-in-use. Amazon’s espoused theory is clear: AI should drive productivity, quality, and customer outcomes. Its theory-in-use, revealed by what gets surveilled and ranked, is something narrower: token volume. When those two diverge, our research and Argyris’s both find the same pattern. Employees calibrate to the theory-in-use because that’s the one with consequences. Cognitive dissonance does the rest. Corporate Leadership Council research has found that a substantial share of departures from companies with strong stated cultures stem from this gap between what leadership says and what its systems actually reward.

Gaming behavior is a predictable response when the measurement system tells employees one thing, and leadership says another.

Thanks for reading Lead from the Front! This post is public, feel free to share the knowledge.

What we’d suggest instead requires harder management, not better dashboards. Outcomes that require judgment to assess, rather than relying on easy proxies. Did this engineer ship faster? Was the code quality higher? Did the customer ticket get resolved with fewer cycles? Token counts are easy because they’re automatic. Outcome assessment is harder because it requires managers to actually evaluate the work, which is the job we’ve quietly stopped expecting managers to do.

There’s also a behavioral dimension that token counts can’t capture. The leading indicator isn’t tokens consumed; it’s whether developers can describe, in their own words, the problem they solved with AI last week. That’s a question a coaching manager can ask in a one-on-one. It can’t be automated, which is exactly the point.

And the leaderboards themselves deserve harder scrutiny. If Amazon’s leadership genuinely doesn’t want token volume gamed, the rankings have to come down. Telling employees the data won’t appear in performance evaluations while continuing to display it on internal leaderboards is a textbook theory-in-use contradiction. Employees see it. They respond accordingly.

The deeper truth tokenmaxxing exposes is one we’ve seen across every major technology adoption cycle. We don’t get the behavior we say we want. We get the behavior we measure and reward. When those two come apart, the measurement always wins.



Response to Ars Technica / Financial Times, May 12, 2026.
Source:
https://arstechnica.com/ai/2026/05/amazon-employees-are-tokenmaxxing-due-to-pressure-to-use-ai-tools/