Skip to content
THE LINUX FOUNDATION PROJECTS
Working Draft June 2026

Production, Consumption, Value

Tokenomics is the discipline forming around a simple chain of conversions: energy and capital become AI tokens, those tokens get consumed to produce intelligence, and that intelligence has to turn into something worth more than it cost to make.

Energy, intelligence, and value. Those three words map onto three buckets, and the buckets are the whole point of this piece, because almost every hard question in AI economics turns out to live in one of them, or in the seams between them.

First, the unit, because the word still trips people. Tokens are the atomic unit of AI, the smallest piece of what every model reads and writes, and it has quietly picked up more jobs than almost any commodity since oil. A token is the output of the data centers and compute we are building. It is the input to how models think. It is what the labs meter and bill against.

Tokens are the thing enterprises are ultimately trying to sell, or sell something built on top of. One unit doing the work of a raw material, a measure, a price, and a product all at once. That overloading is exactly why you need a framework to think about it, and why a single number on an invoice tells you almost nothing.

Riffing off what was presented at FinOps X 2026, thinking about tokens in three conceptual buckets helps frame their behavior like a supply chain.

  • Production is where tokens get made.
  • Consumption is where they get used, and where the real cost is set.
  • Value is where the intelligence they produce either earns its keep or does not.

Each bucket constrains the next, and a bad decision in one shows up as a problem in the others.

How this evolves

This framework is early, and it will change. This is work the Tokenomics Foundation takes on in the open, through vendor-neutral working groups. This is a starting point and a shared language, not a final word.

Production

Production is where the token is manufactured, and its cost is set by physics long before it reaches an API page. For most organizations that means data centers, owned, leased, or simply bought from by the output, and it is starting to spread to the edge as local generation gets cheap enough to matter. The constraint here is real and worth sitting with. The Center for a New American Security calls AI chip production a binding constraint on the pace of the compute buildout, with demand running past what chipmakers forecast and new capacity taking years to stand up. It reaches down to the component level, where high-bandwidth memory has been effectively sold out and advanced-node capacity reportedly several times short of demand. Tokens are scarce because the things that make them are scarce.

What is new in this bucket is sourcing. The same token can come from several places on very different terms, directly from a model provider, through a cloud marketplace, or bundled into someone else’s credit system, each with its own pricing, visibility, and lock-in. That turns production from a make-or-buy decision into a procurement portfolio. Whatever you commit to here, especially the long capacity commitments the current shortage is pushing people toward, sets the floor under every cost downstream of it.

Consumption

Consumption is where tokens get used, and it is where the number that lands on the invoice actually gets decided: who is spending, on what, and whether the spend buys anything. The first surprise in this bucket is that the token bill is the small part. The spend lives in the platform layer wrapped around the tokens, orchestration, agents, memory, retrieval, vector databases, evaluations, governance. A review of AI infrastructure economics puts compute and its surrounding infrastructure at 40 to 60 percent of an organization’s technical budget in the early years of building AI. KV cache is a good illustration: it grows with context length and concurrent users until, on a long-context workload, it can cost more than serving the model itself.

The second surprise is that cheaper tokens have not made anyone’s bill smaller. Per-token prices have cratered; Andreessen Horowitz has tracked inference costs falling on the order of 10x a year since 2021. Almost nobody’s AI spend is going down. That is Jevons paradox in plain sight: make a unit cheaper and people consume so much more of it that the total climbs anyway. Usage goes vertical, the surrounding stack grows to match, and the invoice gets bigger while every individual token gets cheaper. This is the bucket where that fight is won or lost.

The lever that matters most here is efficiency, and it compounds. The teams furthest along treat consumption as a stack, silicon and hardware at the bottom, then capacity, the inference stack, quantization, and model selection, routing, and governance up top. Gains multiply across the layers rather than just adding, so it pays to work the whole stack instead of chasing one optimization in isolation. Routing is the sharpest example. Model prices run from cents to tens of dollars per million tokens, so sending each request to the cheapest model that can actually handle it saves real money. The catch lives inside the same mechanism. Switch models mid-conversation and you can break the prompt cache, and the cheaper model you routed to turns expensive the second it has to rebuild that context from nothing. The upside and the risk are the same machinery.

One caution before leaving this bucket. Most of what sits on an AI dashboard today is probably the wrong thing to be measuring, because the field has not figured out what the right things are. Standardize a metric too early and people start optimizing for the metric, at which point it stops telling you anything. That is an argument for working the measurement out in the open rather than a hundred teams quietly inventing a hundred incompatible numbers.

Value

Value generation requires a return on the investment into the intelligence, and it is the bucket the boardroom now asks about first. What is all this worth, why does it cost what it costs, and is the return there. It is also the least charted of the three, because it reaches past cost management and into the business model itself.

The visible sign is pricing. Software companies are tearing up models they have run for decades and moving off seats and licenses toward consumption. Gartner expects at least 40 percent of enterprise SaaS spend to shift to usage-, agent-, or outcome-based models by 2030, with seat-based revenue shrinking as AI agents start behaving like users that a seat count cannot capture. Andreessen Horowitz has called AI the trigger for a possibly bigger pricing shift than SaaS has been through before. It is arriving first as opaque credit systems, the kind where every action feels like dropping another quarter in the slot, and moving slowly toward direct pass-through. We are going from opaque-but-expensive to clear-but-still-expensive, and what’s interesting is that transparency is improving faster than price.

The reason to treat the three as a supply chain, and not three separate concerns, is that they are wired together. What you pay in production meets the efficiency you build in consumption, and the two of them set your margin in monetization. For a public company that arithmetic can move the stock. Route to the wrong model and break your cache, miss a forecast, sign the wrong capacity commitment, and it does not stay contained in the bucket where it happened; it travels downstream to what your own customers pay. This stopped being a software story a while ago. It is showing up in banks and in any other business now built on top of AI.

What we still do not know

The buckets are a map, not an answer, and the open questions are the real agenda. What do we measure, and against what. How do you standardize pricing across structures that have nothing in common. How do you price the uncertainty, and the experiments that fail, given that plenty of them will. How should token access map to the people doing the work. None of these has a settled answer yet, and anyone selling certainty this early is worth being skeptical of.

A vendor-neutral home for the work

That is what the Tokenomics Foundation is for: a place to work these questions out together rather than a hundred companies solving the same problem in private. It puts the largest consumers of the technology in the same room as its suppliers, and on the supply side that is a new cast, the clouds alongside inference providers, the newer specialized clouds, hardware makers, and the frontier labs.

The work runs through a governing board, a technical steering committee, and IP-managed working groups producing vendor-neutral specifications and best practices. It also has something to build on: the FinOps Open Cost and Usage Specification (FOCUS) already normalizes billing data into a common shape, and extending it to cover AI tokens is the obvious next move toward making the three buckets measurable instead of theoretical.

This is greenfield, and it is bigger than any one company can take on alone. Production, consumption, value is a way to organize the problem, not a claim to have solved it. The work of turning that map into specifications, metrics, and shared practice is the part still ahead, and it is the part worth doing together.