Amazon scored its engineers by AI usage. They beat the game in a week.
June 19, 2026 · Kevin Kelly, Out of Control~6 min read
In a single 30-day stretch, the staff at one Silicon Valley giant fed sixty trillion tokens through their AI tools. Sixty trillion — a number with no human handhold, enough text to redraft the collected works of every author alive several times over. It was not the output of feverish productivity. It was the output of a scoreboard. Somewhere, a leaderboard had whispered to eighty-five thousand people that more tokens meant a better engineer — and the engineers, being human, believed it, gamed it, and built a machine to win it.
The leaderboard that ate itself
Start with Amazon, where the story is cleanest. The company built an internal dashboard called KiroRank that scored developers by their activity on Kiro, its AI coding platform, and on MeshClaw, an in-house agent that can deploy code and answer Slack. The intent was wholesome: Amazon wants more than 80% of its developers using AI weekly, part of a roughly $200 billion 2026 spend on AI infrastructure, and a leaderboard felt like a friendly nudge. Then the nudge curdled. Employees began feeding trivial and fabricated tasks to the agents purely to climb. They had a name for it: "tokenmaxxing." One worker told the Financial Times, "There is just so much pressure to use these tools. Some people are just using MeshClaw to maximise their token usage." Token consumption rose. Cloud bills rose. Useful work did not. Amazon SVP Dave Treadwell sent a memo conceding the system was built with "good intentions" and pleading, "Please don't use AI just for the sake of using AI." Then the company killed KiroRank.
This is Goodhart's law, wearing a hoodie
There's a 50-year-old rule for exactly this, and it deserves a picture you can hold. Imagine a factory paid its nail-makers by weight, so they forged a few enormous, useless nails; switch to paying by count, and they churn out thousands of tiny, useless tacks. The moment a measure becomes a target, it stops measuring the thing you cared about. KiroRank wanted to measure productivity. The instant it became the prize, it stopped measuring productivity and started measuring how good you were at feeding the leaderboard. Token count was always a proxy — a stand-in for "this person is getting real value from AI." Reward the proxy, and people optimize the proxy. They are not being lazy or dishonest, mostly; they are doing precisely what you paid them to do. You asked for tokens. You got tokens.
Amazon's KiroRank scored developers by activity on its Kiro AI platform; staff gamed it by feeding trivial tasks to AI agents — 'tokenmaxxing' — inflating token use and cloud costs. SVP Dave Treadwell told staff: don't use AI just for the sake of using AI. Amazon swapped the leaderboard for 'normalised deployments', which measure useful output. Meta saw the same: an employee-built 'Claudeonomics' board ranked 85,000 staff by tokens, surfacing 60 trillion tokens burned in 30 days. Kevin Kelly's Out of Control explains the loop: reward a proxy and a complex system optimizes the proxy, not the goal. Popular tech commentary; figures per original reports.
Kevin Kelly: you can't command a complex system
Decades ago, in Out of Control, Kevin Kelly spent a whole book on a single uncomfortable truth: the interesting systems of our age — markets, ecosystems, the internet, a company of 85,000 people — are not machines you operate but living systems you can only tend. A clock you command: turn the key, the hands move. A garden, a beehive, a market, you cannot. You can't order a beehive to make more honey, or a forest to grow faster, or a workforce to be more productive. What you can do is change the conditions — and then watch, with a mix of awe and alarm, what the system grows in response. Kelly's phrase for the only kind of control that works on living systems is telling: you give the system its freedom, set the incentives, and it self-organizes into behavior you did not design and could not have predicted.
A corporate leaderboard is exactly such a living system. Management did not type "go run junk tasks" into anyone's terminal; they couldn't have. They set one incentive — rank by tokens — and eighty-five thousand semi-autonomous agents, each optimizing their own standing, self-organized into a token-burning machine. That emergent machine was nobody's plan and everybody's doing. This is Kelly's deepest point, and it stings: in a complex adaptive system, you do not get the behavior you intend. You get the behavior your incentives reward. The leaderboard never measured productivity. It opened a new game, announced the rules, and the org grew the optimal strategy for that game — which happened to be worthless.
Meta proved it wasn't a fluke
If Amazon were alone, you could call it a bad dashboard. But Meta ran the experiment again, by accident. An employee built a leaderboard nicknamed "Claudeonomics," after Anthropic's model, ranking all 85,000 staff by token use and handing out titles like "Token Legend" and "Cache Wizard." It surfaced that astonishing figure — sixty trillion tokens in thirty days, with one top user alone burning 281 billion. At list prices that's an estimated $180 million-plus a month, by some reckonings closer to $900 million. Mark Zuckerberg didn't crack the top 250. The board came down two days after the press found it. Same incentive, same species, same emergent result: gamify a proxy and the hive optimizes the proxy. Two of the most sophisticated engineering organizations on earth, each independently discovering that you cannot bolt a KPI onto a living system and expect it to stay measuring what you meant.
What this means for what you measure
So if you run a team, or a product, or just your own habits, the lesson isn't "never measure." It's that every metric you reward is an instruction to a living system, and it will be obeyed too literally and too well. Before you put a number on a wall, ask Kelly's question: am I treating this like a clock I can command, or a garden I can only tend? Then run the cheap test — if someone wanted to max this metric while producing nothing of value, could they, and would they? Tokens fail it instantly. Notice what Amazon reached for instead: "normalised deployments," a measure of whether AI actually shipped useful code. It's harder to game because it's closer to the real goal — not "did you use the tool" but "did the tool produce something worth using." Measure outcomes, not motion. And hold the metric loosely, because the moment it becomes the prize, your people will, with perfect rationality, grow a strategy you never wanted — and you will have, quite literally, no control.
You can't command a complex system — you set the incentives, and it grows the behavior.
The leaderboard never measured productivity. It announced a new game, and 85,000 people grew the optimal strategy for it — which turned out to be worth nothing.
Framework drawn from Kevin Kelly's Out of Control (凯文·凯利《失控》) — complex adaptive systems, self-organization, and why you tend a living system rather than command it; the metric-as-target failure is Goodhart's law. Facts from May–June 2026 reporting on Amazon's KiroRank and "tokenmaxxing" (Financial Times via The Decoder, India Today, HCAMag; SVP Dave Treadwell's memo) and Meta's employee-built "Claudeonomics" board (The Information via Fortune; 60 trillion tokens / 30 days per The Pragmatic Engineer). Popular tech commentary; figures per the original reports.
技术
亚马逊给工程师按 AI 用量打分,员工一周就把这局玩穿了
2026 年 6 月 19 日 · 凯文·凯利《失控》约 5 分钟
在短短 30 天里,一家硅谷巨头的员工往自家 AI 工具里灌进了六十万亿个 token。六十万亿——一个人类直觉根本抓不住的数字,足以把当今在世所有作者的全集重写好几遍。这不是疯狂生产力的产物,而是一块记分牌的产物。某处,一张排行榜对八万五千个人耳语:token 越多,工程师越强——而工程师,身为人,信了它、钻了它,还造了一台机器去赢它。
那张把自己吃掉的排行榜
先看亚马逊,这个故事在它身上最干净。公司搭了一个内部看板叫 KiroRank,按开发者在 Kiro(它的 AI 编程平台)和 MeshClaw(一个能部署代码、还能回 Slack 的内部智能体)上的活跃度打分。初衷很正:亚马逊希望 80% 以上的开发者每周都用 AI,这是它 2026 年约两千亿美元 AI 基础设施投入的一部分,一张排行榜看起来像个善意的轻推。然后这一推就馊了。员工开始拿琐碎甚至编造出来的任务去喂智能体,纯粹为了往上爬。他们给这事起了个名字:「tokenmaxxing」(刷量)。一位员工对《金融时报》说:「用这些工具的压力实在太大了。有些人就是在用 MeshClaw 把自己的 token 用量拉满。」token 消耗涨了,云账单涨了,有用的活没涨。亚马逊 SVP 戴夫·特雷德韦尔发了封备忘录,承认这套系统「出发点是好的」,并恳求:「请别为了用 AI 而用 AI。」然后,公司把 KiroRank 砍了。
这是穿了卫衣的古德哈特定律
对这种事,有一条 50 年的老规矩,而它值得一个你能攥在手里的画面。想象一家工厂按重量给造钉子的工人计酬,于是他们打出几根又大又没用的巨钉;改成按数量算,他们就哗哗吐出成千上万根又小又没用的图钉。一个度量一旦变成目标,它就不再度量你当初在乎的那个东西。KiroRank 想度量的是生产力。可它一变成奖品,就不再度量生产力,转而度量「你有多擅长喂这张排行榜」。token 数从头到尾都只是个代理指标——它替「这个人正从 AI 里拿到真价值」站岗。你奖励代理,人就优化代理。他们多半不是偷懒、也不是不诚实;他们只是在精确地做你花钱让他们做的事。你要的是 token,你就得到了 token。
亚马逊的 KiroRank 按开发者在 Kiro AI 平台上的活跃度打分;员工拿琐碎任务喂 AI 智能体来刷榜——即「tokenmaxxing/刷量」——推高 token 用量与云成本。SVP 戴夫·特雷德韦尔对员工说:别为了用 AI 而用 AI。亚马逊改用「normalised deployments」(是否产出有用代码)衡量。Meta 也一样:员工自建的「Claudeonomics」榜给 85000 人按 token 排名,曝出 30 天烧掉 60 万亿 token。凯文·凯利《失控》点破这个回路:奖励一个代理指标,复杂系统就去优化那个代理、而非目标。大众科技评论;数据以原报道为准。
所以如果你带一支队伍、做一个产品,哪怕只是管自己的习惯,这一课不是「别度量」。而是:你奖励的每一个指标,都是给一个活系统下的指令,它会被照办——办得太字面、太到位。在你把一个数字挂上墙之前,先问凯利那个问题:我是在拿它当一座我能命令的钟,还是当一座我只能照料的花园?然后跑那个便宜的测试——要是有人想把这个指标拉满、却什么有价值的东西都不产出,他做得到吗?他会去做吗?token 这一关瞬间就挂。注意亚马逊转而抓住了什么:「normalised deployments」,一个衡量 AI 到底有没有交付出有用代码的指标。它更难刷,因为它离真正的目标更近——量的不是「你用没用工具」,而是「这工具有没有产出值得一用的东西」。量结果,别量动作。而且对指标要松松地拿着,因为它一旦变成奖品,你的人就会以完美的理性,长出一套你压根没想要的策略——而那时你,确确实实地,毫无控制可言。
框架取自凯文·凯利《失控》(Kevin Kelly, Out of Control)——复杂自适应系统、自组织,以及「活系统只能照料、无法命令」;「度量变目标即失真」即古德哈特定律。事实来自 2026 年 5–6 月对亚马逊 KiroRank 与「tokenmaxxing」的报道(《金融时报》经 The Decoder、India Today、HCAMag;SVP 戴夫·特雷德韦尔的备忘录),及 Meta 员工自建的「Claudeonomics」榜(The Information 经 Fortune;30 天六十万亿 token 据 The Pragmatic Engineer)。大众科技评论;数据以原报道为准。
枠組みはケヴィン・ケリー『コントロールの喪失』(Kevin Kelly, Out of Control)より——複雑適応系、自己組織化、そして「生き物は命令でなく世話するもの」。「尺度が目標になると壊れる」はグッドハートの法則。事実は2026年5〜6月のアマゾン KiroRank と「tokenmaxxing」の報道(フィナンシャル・タイムズ経由 The Decoder、India Today、HCAMag;SVP デイブ・トレッドウェルのメモ)、および Meta 社員製「Claudeonomics」盤(The Information 経由 Fortune;30日で60兆トークンは The Pragmatic Engineer による)。大衆向けテック評論。数値は原報道を基準に。