Skip to content
THE LINUX FOUNDATION PROJECTS

Tokenomics Foundation · Technical Steering Committee

What ten intro calls tell us, and what the working groups should build first

A synthesis of the pre-kickoff one-to-one conversations with incoming TSC members, and a proposed opening backlog for the four working groups running the first four-week sprint into Amsterdam.

10 transcripts reviewed in full 4 working groups 4-week sprint Artifacts due before Tokenomicon+X, Amsterdam, 22–23 Sept
On attribution. Nothing here is attributed to a person or a company. Positions are reported by how many independent conversations raised them, so the group can weigh an idea on its merits rather than on who said it. Source summaries with full attribution live in the vault alongside the transcripts.

01Where the room already agrees

Eleven positions surfaced in more than one conversation. The count is the number of separate calls in which the idea was raised unprompted or affirmed strongly. Anything at seven or above should be treated as settled enough to build on rather than debate.

What each of those actually means

Tokenomics is not token counting 9 of 10

Every substantive call independently rejected the narrow reading of the name. The framings varied -- unit economics, ROI, cost to serve and margin, an operations discipline in the lineage of DevOps and FinOps, energy per completion -- but nobody arrived defending token counting as the scope. One participant put it flatly: tokens are the focus, but counting them is not the work. This is the single strongest signal in the set and the definitions group should be able to ratify it in week one without debate.

Nobody agrees on the denominator 7 of 10

The recurring question is not what tokens cost but what to divide that cost by. Candidates raised across calls: the commit, the pull request, the resolved support ticket, the completed task, the remote session, the intent, the feature, the seat. One participant counted roughly ten plausible options and no shared account of the tradeoffs. A second observed that cost per completed task sounds excellent on a slide and collapses the moment you ask what "completed" means. This is the central open problem of the Value group.

A token is not a fungible unit 6 of 10

Cached, standard, priority, output, reasoning and fine-tuning tokens carry different prices and different physical costs -- one participant put the spread at roughly twentyfold within a single model. Across providers the tokenizer and serving stack differ, so the same text is a different number of tokens. One framing compared model choice to a CPU architecture migration rather than a price-per-unit optimisation. The implication is uncomfortable: any efficiency metric built on a blended token rate is measuring the wrong thing.

A FOCUS-shaped standard is the most-wanted artifact 6 of 10

Six calls independently named FOCUS as the model to copy, though they disagreed on where to point it: at model provider billing, at SaaS vendor invoices, at runtime telemetry, or at a cross-provider metadata and tagging convention. Notably, more than one person valued FOCUS less for its schema than for the shared vocabulary it forced. That is a hint about what a v1 should optimise for.

The real cost sits below the token layer 6 of 10

Vector stores, retrieval and embedding pipelines, warehouse scans fired by agents, orchestration compute, storage of inputs and outputs, tool-call side effects. One participant reported that across client deep-dives more than half of application cost sat below the token line in some cases, meaning token-only ROI reporting is wrong by roughly half. Another noted that once an agent pushes enough data through a warehouse, tokens become the negligible line item. A third tracks nine distinct AI cost categories, of which tokens are one.

Telemetry, not billing, is the substrate 5 of 10

The bill will never carry the AI story. The standards that matter are what applications emit at runtime, and today that layer is fragmented across OpenTelemetry GenAI conventions, vendor-specific formats, gateway logs and hand-rolled proxies and hooks. Several participants are already reverse-engineering it per customer. This reframes the Foundation's first technical deliverable from a normalisation spec to an instrumentation spec -- and, per one proposal, from a document to shipped code.

Attribution is broken, and tagging will not save it 5 of 10

Tagging worked when spend mapped to owned resources. In a consumption world the resource is not owned, and vendors have moved functionality into their own clouds where tags do not propagate. Provider-side allocation stops at the inference boundary and does not follow the downstream function invocation or file write. In agent-to-agent flows a product owner cannot separate their own consumption from consumption another agent triggered. The workaround everyone has landed on -- proliferating accounts, projects and API keys -- is already unmanageable.

Definitions must go first, and be time-boxed 5 of 10

Several people said explicitly that the definitions work gates everything else and should be finished, not perfected. The most concrete version of the ask is a single framework diagram in the shape of the FinOps Framework, because that picture is what gets pasted into an internal strategy deck and what makes a first-time viewer understand the field. Sales conversations are currently failing on vocabulary alone: a customer says they want to talk about AI and nobody can route the conversation.

Agentic spend is a live governance risk 4 of 10

Fewer calls raised it, but with the highest intensity of anything in the set. Buggy agents loop in production and still earn recognition because engineers are incentivised on shipping. Anyone with chat access can start generating spend outside every existing limit. The named fear is a runaway agent producing an overnight anomaly large enough to become a reputational story rather than a cost story. One participant reported an organisation that stood up a token-usage leaderboard and discovered months later that engineers had built self-replicating agents to climb it.

Multi-provider is the day-one default 4 of 10

In cloud, multi-provider was an advanced maturity scenario. Here it is where everyone starts, mixing hyperscaler-hosted models, first-party APIs, neoclouds, open-weight models and self-hosted inference from the outset. Normalisation therefore cannot be deferred to a later maturity stage the way it was in FinOps.

Ship written artifacts fast 4 of 10

The clearest process advice in the set: prioritise what the group can decide unilaterally, without waiting for model providers to change anything, and publish while attention is high. Circulate written drafts for comment rather than seeking consensus in meetings, because asking for a minute in a call reliably costs five. One participant argued that much of the FinOps framework's growth was obvious in hindsight and only needed someone willing to write it down.

02Outlier ideas worth putting in front of the group

Each of these came from a single conversation. None of them are fringe -- they are load-bearing observations that the rest of the group has not yet heard, and several would change what a working group builds first. Chips indicate the group most likely to own the idea.

Demand and forecasting is a missing pillar

The current map has production, consumption and value, with nothing for demand. Forecasting is what makes procurement possible, and one organisation described a customer with a full-time person dedicated solely to buying tokens who said outright it was not sustainable. Treat it as a fifth domain or an explicit cross-cutting thread.

DefinitionsProduction

Energy belongs at the consumption layer too

Energy is currently scoped to production. But the consumption side determines how much energy a workload draws, and the practitioner has far more control there than over the data centre. The overlap between what raises cost and what raises electricity is tighter in AI than it ever was in cloud.

ConsumptionProduction

Carbon is an accounting exercise; electricity and water are physics

Carbon is a coefficient bolted onto energy and can be moved tenfold by accounting choices. Electricity and water are the actual limits on capacity. A framework that leads with carbon inherits the gaming problem; one that leads with energy does not.

ProductionDefinitions

Call it a premium, not an inefficiency

The gap between reserved-capacity and pay-as-you-go pricing is not waste; it buys availability, capacity and SLAs, the way an enterprise database costs more per gigabyte than the SSD under your desk. One team computes a premium index across models and deployments. It reframes a whole category of "waste" findings as deliberate purchases.

ValueConsumption

Quality is the second premium axis

Extending the same logic: you now pay a premium for a better model to get better output, just as you once paid for faster page loads. The difference from cloud is that degradation used to mean latency on the same page; now the literal words returned to the customer change. Availability premium and quality premium combine into one measurement frame.

Value

Outcome-based pricing is partly a transparency defence

The commercially uncomfortable observation in the set: some vendors are pushing outcome-based pricing precisely to avoid token transparency, because the moment you sell tokens the customer points out they could buy them elsewhere for a tenth. Some of the vendors the Foundation needs may actively resist the standards it wants to publish.

Value

Goodhart's Law is the governing risk

When a measure becomes a target it stops being a good measure. The obvious first artifact -- a token usage leaderboard -- has already failed in the field. A standards body that publishes metrics without publishing their gaming failure modes will make the problem worse, not better.

DefinitionsValue

Rework economics can invert the cheaper model

Dropping from 99% to 96% accuracy across two million documents means tens of thousands of failures and human remediation that erases the model saving. Model routing and right-sizing are a red herring wherever accuracy matters, and you cannot know a model's real accuracy until roughly ten thousand sandbox runs. A cheaper model can also look better on cost per pull request precisely because it ships bugs that need more pull requests.

ConsumptionValue

Input tokens are the problem, not output tokens

Output length is largely uncontrollable -- a model can simply be chattier, reasoning modes multiply it, prompt-level caps get exceeded. Input is where organisations push hundreds of billions of tokens, degrade accuracy, and pay twice. This inverts where most optimisation guidance points.

Consumption

Make providers disclose the tokenizer

A token without provenance is as meaningless as a gram of carbon without a methodology. If billing files disclosed the tokenizer and serving stack, tokens would become comparable across providers. This is a concrete, narrow ask the group could put collective weight behind.

DefinitionsProduction

Standardise telemetry, and ship it as code

Rather than a schema for what providers hand you after the fact, a standard for what applications emit at runtime, aligned to OpenTelemetry conventions and delivered as an SDK with drop-in snippets. Ten agentic applications in one enterprise would then emit one consistent stream. It relocates standard-setting from the billing seam to the instrumentation seam -- and it would be the Foundation's first code deliverable.

Consumption

Ship dashboards, not just specs

A parallel proposal from a different call: the fastest path to practitioner adoption is an open-source library of pre-built dashboards and queries for token unit economics, in the spirit of the well-known AWS dashboard project, hosted by the Foundation. Practitioners adopt what they can deploy on a Friday.

ConsumptionValue

Human development effort belongs in the denominator

To do fair value you have to account for everything the token touches, and probably the engineering time that built the thing, which never appears on a provider bill. It is a scope expansion few will raise and it points straight at capitalisation and accounting treatment.

Value

Maybe the Foundation should not standardise value

A deliberately contrarian position, and notable for coming from a vendor who would benefit from the opposite. Value is whatever a company puts at the core of its business, so it is the least standardisable layer. The leverage is upstream: standardise the data and the vocabulary, and let each organisation define value on top. Worth an explicit decision rather than an assumption.

ValueDefinitions

Ship a decision system, not one metric

The complementary position: one unit of value will not fit everyone, so enumerate the candidate denominators and document the tradeoffs of each rather than ratifying a single number. A comparative framework, not a prescription.

Value

Reduce your model count to fix forecasting

Contrarian against the routing consensus. One organisation was found running fourteen models across three platforms. Fewer models means predictable consumption, which means forecastable demand, which means procurement becomes possible. Sits directly against the "route dynamically" instinct and deserves an argued position.

ConsumptionProduction

A token factory need not own metal

One participant brokers spare capacity from other GPU providers and surfaces it as ordinary nodes in a customer's cluster. That decouples "token factory" from "owns hardware" and complicates any production framework that assumes the producer controls energy and procurement.

Production

The atomic unit for agents is the tool call

Agent platforms deserve their own taxonomy branch, peer to compute and storage, with tool invocations, agent counts and licence distribution as first-class metrics. Most FinOps-derived framings stop at the model and the token, which is one level too high for agentic workloads.

DefinitionsConsumption

Optimise at three layers, not one

Model, inference, and application harness. Most discussion collapses to model choice and prompt hygiene, but the harness -- the agent loop itself -- is where a lot of the waste is generated, and it is the layer an organisation building its own agents fully controls.

Consumption

Optimise on the endpoint, before the request leaves

A small model resident on the developer's machine, hooked into the coding agent, compressing prompts and classifying the type of work requested before anything reaches a provider. The field assumes the optimisation surface is the API call or the model choice; this moves it upstream of both.

Consumption

Treat waste like a security detection catalog

Framing cost work as detection and response rather than reporting suggests an artifact shape nobody has proposed: a shared, versioned catalog of AI waste detections expressed as signatures, contributed and maintained the way threat signatures are.

Consumption

Small local models are coming, and that changes the question

The argument that open-weight quality is already good enough and that specialised models will increasingly run locally. If that holds, tokenomics is less about negotiating rates with frontier vendors and more about a recurring build-versus-buy inference decision -- a very different framework.

ProductionConsumption

Infrastructure as prompts

Everyone in the organisation now effectively has console access. The stakeholder surface for whoever owns tokenomics has grown by orders of magnitude, reaching procurement, sustainability, new graduates and every line of business. That is a reach opportunity and a governance problem in the same observation.

Definitions

Sustainability gets worse as you move closer to the metal

Moving from managed inference to raw GPU and bare metal materially degrades sustainability posture, and it is unresolved which emissions scope token consumption even lands in. A direct input to whether energy and sustainability sit inside production or stand on their own.

Production

One primary actor per domain

Rather than a flat persona list across all of tokenomics, give each domain a single accountable owner with everyone else in support -- procurement owns supply, and so on. A cleaner accountability model, and a better fit for how the work actually lands in an enterprise.

Definitions

Hard value brings soft value; soft value does not reciprocate

Hard value is revenue up, cost down, labour down. Soft value is satisfaction and performance. The asymmetry is a usable rule for sequencing a value framework: lead with the hard measures and the soft ones follow, not the other way round.

Value

Stateless tokens make a spot market conceivable

Token generation is stateless in a way cloud workloads never were, so a request can be served anywhere. Layer latency tolerance onto routing and you get something like spot pricing for inference, up to and including a facility that powers down half its capacity and reroutes when the grid needs it back.

Production

03Live disagreements to surface, not settle quietly

Four genuine conflicts appeared across the calls. Each is currently invisible to the group because the positions were stated in separate rooms. Putting them on an agenda is more valuable than resolving them by default.

Should the Foundation standardise value at all?

One position holds that value is inherently company-specific and the Foundation's leverage is entirely upstream, in data and vocabulary. Others want an explicit value and ROI methodology as the headline deliverable, and one asked for prescriptive guidance from an authoritative neutral voice. There is a Value working group either way, so the group should decide whether its output is a method, a decision framework, or a standard.

Route dynamically, or consolidate models?

Three separate positions. Routing as the central optimisation lever, with metadata steering workloads to the right model. Consolidation, because fewer models makes consumption forecastable and procurement possible. And routing as a red herring wherever accuracy matters, because rework costs more than the model saved. All three are defensible; they imply different artifacts.

How far does the scope run?

Consensus rejects token counting, but there is real unease about the other end. One participant was openly surprised that the scope is effectively all of AI cost, because customers hear the name and think token management. Another reported analyst pushback that the name is simply wrong for the scope. Meanwhile the sub-token evidence pushes scope outward. The definitions group needs to draw and defend both boundaries.

Where does energy live?

Currently scoped to production, but the strongest energy voice in the set works entirely at the consumption layer, and a separate participant wants power consumption tied to business value on the production side. There is an unresolved question above both: whether sustainability is a thread inside production or a group of its own.

04Proposed opening backlog

Two to three items per group for the four-week sprint, ordered so that the first item is the one to start on day one. Every item is scoped to be finishable and publishable before Amsterdam, and deliberately chosen to need nothing from model providers. Cross-cutting note: the definitions vocabulary gates the other three groups, so it should ship in draft by end of week two rather than end of sprint.

Definitions, Personas & Frameworks

Tuesdays, 08:00 PT

In-flight: Personas + Operating Model Map (draft ~90%); Tokenomics 101 content.

1

Tokenomics Framework v0.1 -- one page, publishable

A single diagram in the shape of the FinOps Framework, with the domains, the lifecycle phases and the personas on one canvas. Ship it deliberately incomplete and dated. Time-box hard: two weeks to a circulated draft, four to publication.

Why first: the most concrete artifact request across the whole set, and the thing members need to paste into internal decks to justify participation. Nine of ten calls implied it; several asked for it by name.

2

Controlled vocabulary v0.1

Three tables. Token classes (input, cached input, output, reasoning, fine-tuning) with what makes each non-comparable. Workload archetypes: conversational assistant, internally-built agent or workflow, and customer-facing agentic workflow -- each with the granularity and success measure it needs. And the production / consumption / value boundary written as prose, including where demand and forecasting sit and whether TokenOps is a synonym or a layer.

Why second: the fungibility problem and the missing archetype taxonomy each surfaced repeatedly, and commercial conversations are reportedly failing on vocabulary alone. This is also the dependency the other three groups are waiting on.

3

Persona map with one primary actor per domain

Name the accountable owner for each domain and put everyone else in support. Include procurement and sustainability alongside the engineer, product manager and finance roles, and account for the fact that effectively everyone in the organisation now generates spend.

Why third: smaller, and it depends on the domain boundaries settling first. Two calls made a specific case for it and one flagged the sustainability persona as consistently missing from comparable frameworks.

Production: Token Factories & Energy

Wednesdays, 13:00 PT
1

Reference model of the token factory stack

Land and power commitment, hardware, serving stack, quantisation and model configuration, up to the served token -- with the cost driver and the controllable levers named at each layer. Must accommodate the producer who owns no hardware and brokers spare capacity, because that case breaks the obvious assumptions.

Why first: production is the least understood domain among the members themselves -- more than one asked what it even covers. A layer model is the cheapest way to make the domain legible and to recruit into it.

2

Energy metric set and the production/consumption boundary

Define energy per token and energy per completion, PUE, and utilisation, then resolve explicitly which of these belong to production and which to consumption. Lead with electricity and water as physical constraints and treat carbon as a derived, methodology-dependent coefficient. State the emissions-scope question for token consumption even if the answer is deferred.

Why second: this is the group's clearest live disagreement, it was flagged as a genuine gap in Foundation expertise, and it is a decision the group can make on its own authority.

3

Self-hosted versus API: a worked cost model

One fixed workload, costed both ways, showing where the crossover sits and what the model omits -- utilisation risk, engineering time, quality delta, and the sustainability penalty of moving closer to the metal. Cover the open-weight substitution case explicitly.

Why third: multiple organisations are already running mixed in-house and open-weight estates and making this call without a shared model. It is also the most directly useful artifact for anyone weighing local inference.

Consumption: Efficiency & Optimization

Thursdays, 10:00 PT

In-flight: Big-T Notation (multiple pieces recently released for discussion); Optimization Playbook (draft ~50%).

1

AI telemetry schema v0.1, aligned to OpenTelemetry

What an AI application should emit at runtime to make cost attributable: model and token classes, but also the sub-token layer -- retrieval and embedding calls, vector store and warehouse queries, tool invocations, agent identity and parent agent. Publish as a spec plus a reference implementation and drop-in snippets in a Foundation repository, so the deliverable is code as well as a document.

Why first: the single most-requested technical artifact, and the dependency underneath almost every other ask. It is also entirely within the group's control -- it requires nothing from model providers. The Foundation being chartered for software projects makes this the natural first proof of that.

2

Waste and anti-pattern catalog v0.1

Named, detectable patterns with the signal that identifies each and the correction: input context bloat, reasoning modes on tasks that do not need them, agentic rollouts on trivial work, hard-coded model pinning across agent fleets, cache misconfiguration, idle licences, looping agents. For each, state the detection signature and the expected saving. Include the counter-case: where the cheaper path costs more once rework is counted.

Why second: immediately usable by practitioners, it is where the specific field observations concentrate, and framing it as a signature catalog makes it a living artifact the community can extend rather than a one-off document.

3

Agent spend guardrails and the post-tagging attribution model

Budgets, rate limits, kill switches, anomaly thresholds and escalation paths for autonomous agent spend, plus a documented attribution approach for a world where tagging does not work -- including agent-to-agent flows where consumption crosses ownership boundaries. Note explicitly where each metric can be gamed.

Why third: the highest-intensity fear in the interviews, and one that is currently being solved privately and badly. It depends partly on the telemetry schema, so it starts once that draft exists.

Value: Financial Reporting, Capitalization & Accounting Treatments

Fridays, 09:00 PT

In-flight: Unit Economics and Business Measurement (draft ~75%).

1

Unit-of-value decision framework

Enumerate the candidate denominators -- commit, pull request, resolved ticket, completed task, session, intent, feature, seat -- and for each give the workload type it suits, the data required, the gaming failure mode, and what it systematically hides. Ship it as a decision aid rather than a single ratified metric, since the evidence says one number will not fit everyone.

Why first: the most-cited open problem across the whole set, and the pluralist framing resolves the standing tension between wanting a standard and value being company-specific. Deliverable, publishable, and needs nothing from anyone outside the group.

2

Completion framework

Not a definition of "complete" -- that is organisation-specific -- but a repeatable method for an organisation to arrive at its own: how to identify task types, who needs to be in the room, how to set the bar, and how to handle accuracy and rework so the correction is inside the number. Address the self-certification problem, since a model judging its own completion is not evidence.

Why second: cost per completed task is the metric everyone reaches for and nobody can compute. A method that makes it computable is the highest-leverage thing this group can produce, and it was explicitly volunteered into this group.

3

Cost-to-serve and the premium index

What belongs in the denominator: sub-token infrastructure, orchestration, and engineering time that never appears on a bill. Then the premium model -- separating deliberate purchases of availability and quality from genuine waste -- and the capitalisation and accounting treatment of model development, fine-tuning and inference spend. Stretch item: what an AI-enabled software vendor's outbound invoice should disclose when it is not selling tokens directly.

Why third: the accounting content is the group's charter and its most defensible territory, but it depends on the denominator work landing first. The premium reframing is the most transplantable single idea in the interviews.

05Two things to decide at kickoff

Not backlog items -- blockers that will otherwise be resolved by default.

Where does demand and forecasting live?

It spans all three current domains and is owned by none. Either name it a fifth domain or assign it explicitly as a cross-cutting thread with a named owner before the sprint starts.

What shape should the Value group's output take?

Options range from prescriptive standards to methods, guidance, best practices, and case studies. The backlog above assumes a method and decision framework -- a decision aid rather than a ratified metric -- since evidence says one number will not fit every organisation. If the group intends something more prescriptive, the first backlog item changes materially. Decide it now rather than discovering the answer in week three.