Skip to content
THE LINUX FOUNDATION PROJECTS
Insight

Podcast #15 of Tokenomics Brief: Tokenomics Goes Official: Defining AI Economics, Value, Routing Layers & the Per-Watt Economy

J.R Storment

J.R. Storment

KEY INSIGHT

Tokenomics is not about counting tokens

Only about a quarter of a platform’s AI spend is typically direct model consumption. The rest is GPU, infrastructure, memory, database, compute and the labour around all of it — so a cost model that stops at the token line item misses roughly three quarters of the bill, and the value formula built on top of it is wrong.

Listen on Spotify or watch on YouTube.

Summary

The Tokenomics Foundation has moved from announcement to institution. The governing board held its first meeting, the technical steering committee stood up behind it, and the value, consumption and definitions working groups all met for the first time, with production following a day later. The board meets quarterly, works by consensus and owns vision, strategy and budget. The technical steering committee owns ratifying the technical work, including the definition of tokenomics itself.

That definition is staying in draft on purpose. J.R. Storment expects it to change “50 times over the coming years” and treats that as a feature: take strong positions early, in the open, and keep revising as the industry teaches you. What the founding-member onboarding calls produced was less a definition than a diagnosis — the parable of the blind men and the elephant. Hardware people described tokens per watt, finance people described attribution and unit costs, engineers described caching and routing, product people described pricing and margins, and CFOs tried to connect it all on the balance sheet. Everyone was touching something real. Nobody described the whole animal.

The closest thing to consensus was relating cost to outputs, and there was one near-unanimous exclusion: tokenomics is not about counting tokens. The number behind that stance is the sharpest figure in the episode. An audit of those conversations found that only about a quarter of a platform’s AI spend is typically direct model consumption. Stop at the token line item and you are looking at the tip of the iceberg.

The value working group went straight at the hardest question — measuring “before AI” rather than “after AI”, summing both sides, and counting labour on both sides as augmentation rather than headcount reduction. It agreed a set of value categories ranked along a quantifiability gradient: revenue enablement, cost avoidance, capacity gain, labour offsets, output quality, speed to market, risk reduction, and net new capabilities. It also produced the first complete measurement method the foundation has seen — deflection — and immediately found the hole in it. Processes that did not exist before AI have no baseline and no counterfactual, so a framework that only handles deflection looks complete while missing the harder half.

The consumption group has already shipped drafts: the five-layer tokenomics stack and Adobe’s Big T notation, a riff on Big O applied to token consumption complexity. The definitions group spent its time on energy, and that is where the episode lands. Every supply-side conversation Storment has had — neoclouds, data centre providers, GPU, CPU and TPU providers — is converging on metrics with “watt” in them: revenue per watt, token output per watt, and intelligence per watt as proposed by the Agentic AI Foundation. FinOps was built on the premise that capacity could always be procured. Tokenomics cannot assume that, because power runs out first.

Three market stories close the episode, and they read as one pattern: the layers that decide where a request goes, and what power it consumes, are becoming financial infrastructure. Stripe acquired OpenRouter for $7 billion, having already bought metering and billing. SpaceX closed its Cursor acquisition, putting a consumption layer in front of gigawatt-scale compute. And a SpaceX earnings report split out its AI segment and quoted a specific revenue-per-watt figure — the metric leaving the working group and entering public financial reporting. MIT and NYU joined as academic members, and Harvard Business School’s AI Institute landed on a three-word playbook: route, pilot, govern.

Key takeaways

  • Every value claim needs a named baseline. Before AI versus after AI, both sides summed, humans included on both sides. If an AI story cannot name its counterfactual, it is not an ROI story yet.
  • The routing layer is financial engineering infrastructure now. A payments company paying $7 billion for a token router, on top of buying metering and billing, makes token cost management a first-class financial function. Plan the architecture as though the router is a critical cost centre.
  • Cost per call beats cost per token. Token prices do not tell an operator what a single processed invoice or auto-remediated alert costs, and they ignore the human time deflected.
  • Labour belongs in the cost equation, as capacity rather than headcount. The room was careful with the language: augmentation, not cuts.
  • Deflection has a baseline problem. For mature processes the right comparison is the marginal cost of already-optimised, sometimes offshored work, not a fully loaded salary. And once agents absorb the easy cases, only hard ones reach humans and throughput metrics invert.
  • There is a quality floor. A model that is 90% cheaper but misses the quality bar a workload requires is not cheaper. It is unusable for that workload.
  • The watt is becoming the denominator of the supply side. Revenue per watt, tokens per watt, intelligence per watt. A full episode on the per-watt economy is coming.
  • Nobody confused tokenomics with crypto. Across dozens of member onboardings, not once.

Never miss an episode

The Tokenomics Brief lands every week, breaking down AI economics as the discipline forms. Follow it on Spotify or YouTube and each new episode arrives the day it drops.