Tokeneconomics Foundation | Roadmap input
What Members Asked Us to Build
Every item below was requested or explicitly endorsed in a member onboarding or pre-board conversation. Nothing is invented, and nothing is attributed.
Every item below was asked for, or explicitly endorsed, by someone in a member onboarding or pre-board conversation. Nothing is invented and nothing is attributed. Where a request appeared in only one call but was made forcefully, that is marked, because a single strong signal is not the same thing as consensus.
How to read the signal column
Tier 1. Language and definitions
This tier was the loudest and least contested thing in the entire set. Multiple members named it as literally the first thing to do, unprompted, in separate conversations. Two of them opened with it.
| Deliverable | The pain behind it | Signal |
|---|---|---|
| A definition of tokenomics, agreed collectively and published | Everyone describes it differently. Solution providers are improvising against whatever the customer means, with no reference point to push back on. | Unanimous |
| A definition of a token, including token types (input, output, reasoning, image, text) | Providers tokenize differently; billing data lumps all output tokens together; cross-model comparison is currently invalid. Described as “very baseline type of work” that unlocks everything above it. | Strong |
| Taxonomy and shared lexicon: production versus consumption, retrieval tokens, output, prompts, context, cache terms | Customers use the words without shared meaning, which blocks KPI and metric construction downstream. | Strong |
| Naming and defining the full component cost picture of AI, not just tokens | Direct model consumption can be a minority of actual AI spend. People fixate on the token line and miss the majority. | Strong |
| A bill of materials for a token: first-principles breakdown of what it costs to produce one, and what the production options are | Buyers cannot reason about buy versus generate versus self-host without knowing what they are paying for. | Single |
That’s my first thing. Standardization of definitions and technology layers. And methodologies.
I think first it would be the shared language, before the metrics. The language feeds the metrics.
It’s not unifiably defined between the different models. This is very baseline type of work, but it helps comparability in the AI space a lot.
Sequencing note
Tier 2. Measurement and metrics
| Deliverable | The pain behind it | Signal |
|---|---|---|
| A standard recipe for measuring cost, before touching value at all | No agreed method exists. Members are each inventing one. | Moderate |
| Token efficiency normalization: token density, effective token cost, units of work per token, a token grade or price index | A model with attractive per-token pricing can consume twice the tokens for the same task. Per-token price alone is a misleading metric. Suggested method: one prompt across five models, repeated roughly a hundred times, to derive relative generation rates. | Strong |
| Attribution standards mapping AI and token cost to business capabilities, processes, applications, and internal productivity | Requires stitching metering systems, observability data, and token data together. Described as “super complex” even inside a company with fully centralized AI consumption. | Strong |
| Individual-user-level attribution guidance | FinOps has always attributed to an application, service, or department, never to a person. Developer productivity use cases force the change. | Moderate |
| Cost allocation best practices: chargeback, showback, entitlement | Being figured out ad hoc with each customer because no reference model exists. | Moderate |
| Budgeting and forecasting methods for AI spend | Members are being asked for three-year token spend forecasts with no baseline data. One noted that historical trending plus a percentage, the method used for two decades, simply does not apply. | Strong |
| A vendor-neutral token demand forecasting method: tokens per user per day, per workload, per model baseline, aggregated and projected on adoption curves | Software vendors cannot tell customers how many tokens per user their product needs. Buyers get GPU counts with no way to know if that is enough. | Single |
| Production-side metrics: intelligence per watt, tokens per watt, cost per intelligence, cost of a token per megawatt | Data center and hardware members work in dollars per megawatt; nothing connects that to token output. | Moderate |
| A value and impact metric standard correlating token investment to time saved, cost avoided, revenue, growth, new markets | Members are repeatedly told to measure ROI with no method offered. The “I” in ROI is itself undefined. | Strong |
| Standardized cross-company business outcome metrics for comparable processes | Common processes exist at every company, so they can be standardized and benchmarked in a way that bespoke software cannot. | Single |
| A self-service benchmarking framework, so teams can benchmark their own workloads | Published third-party benchmarks do not answer whether a cheaper model is close enough for your job. Explicitly not the Foundation publishing scores. | Moderate |
I’m tired of people telling me they have to measure ROI and having no ideas for it. I’ve sat in plenty of those meetings. I would like frameworks and baseline ideas of success stories where people have tried to measure ROI, and a vision of how you approach that.
Establishing these types of standards is very important. I call it attribution, so that we can really attribute AI to certain business capabilities, being it a process, an application, or internal productivity.
Tier 3. Data specifications and provider pressure
The recurring pattern here: members want the Foundation to articulate the data required for good decisions, then use that articulation as leverage on providers. That was named explicitly as the proven playbook.
- Extend the cost and usage spec for tokens and AI. Deeper token telemetry, a principal or actor dimension so cost attributes to an individual, and provenance handling for the case where one provider’s model is procured through another platform. Requested and endorsed widely; also named internally as the single highest-conviction long-term outcome. Strong
- A worked sample bill in spec format for frontier model providers, used as an engagement artifact: here is what we think your data would look like, tell us where we are wrong or why you cannot do this. Paired with customer demand signals. Single
- A normalized data format for AI productivity tool data spanning code assistants and agent tools, so productivity data is comparable across vendors. Moderate
- Data reporting standardization across cloud platforms, model platforms, and internal systems, including a data model for how it normalizes into something you can query. Moderate
- Task classifiers: a common set of consumption metadata tags describing why AI is being used, defined in common and pushed to providers. One member is already piloting this directly with a provider and asked the Foundation to standardize the classifier set. Single
- Standardized terms and conditions across model providers. Legal review currently consumes “an army of legal people for months and years.” Single
- A standardized credit model. Each provider’s credit system is incomparable, and moving between them is punishing. Single
- An observability and telemetry data model for AI infrastructure, handling high cardinality and multiple reporting depths. Billing tells you a user spent millions of input tokens; telemetry tells you what they were doing. Single
- Agent and hardware asset data, either inside the cost spec or as a defined intersection with IT asset management. Unresolved which. Single
- Cost and usage spec coverage for data center and AI equipment, reviving capacity, energy, and cooling disciplines that were largely retired during the move to cloud. Single
- A required feature set baseline for tools in the space, so buyers know what to expect from a platform. Single
If tokenomics accomplishes little else over time, we’ll be successful, it’d be worth it, if it helps standardize token cost reporting, AI cost reporting.
If we even had some common thoughts on what classifiers could be, and shared that and influenced the providers, I think that could be interesting.
Tier 4. Frameworks, operating models, personas
| Deliverable | The pain behind it | Signal |
|---|---|---|
| A prescriptive framework for managing AI cost, in the style of the FinOps framework, published early and refined in public | Members have no guiding principles at all and want a starting prescription, explicitly accepting that v1 will be wrong. Named as the top ask in more than one call. | Strong |
| A maturity model with domains and subcomponents (visibility, financial control, optimization) and a stated “what good looks like” | Everything is reactive firefighting. Clients have no target state to aim at. | Moderate |
| Personas and operating model: who does tokenomics, how it integrates with FinOps and the rest of the business, how to staff it | Advanced companies have a hybrid principal engineer who can speak both languages. Most companies do not, and need to know how to interface instead. | Strong |
| A framework for building a tokenomics practice from zero | There is no equivalent of the FinOps practice-building guidance. | Moderate |
| A bridge document for FinOps professionals newly assigned AI spend: what transfers, what to tweak, what to discard | They are being handed the problem and guessing. “How do I tweak this, or is this relevant, or am I thinking about this correctly?” | Moderate |
| A “what questions should I ask” guide for the interface between finance-side practitioners and deep technical owners | Practitioners hit a hard technical ceiling and do not know what to ask past it. | Single |
| Canonical architecture diagrams of the AI stack, with every named technique placed on the diagram | The strongest-stated pain in the internal conversations. AI is “an amorphous blob” with no shared mental picture equivalent to a request hitting a load balancer. Readers cannot tell whether a buzzword is their job. | Moderate |
| A landscape and ecosystem map, categories first and then the projects and vendors that fill them | Practitioners cannot locate a tool or project in a category, and confuse similarly named products that sit in completely different layers. | Moderate |
| An AI optimization stack: what happens at hardware, model selection, prompt, context, routing, and caching, and which lever to pull when | No map of where efficiency work actually lives. | Strong |
| Model selection guidance and a model garden, including new versus old versions and open weights versus proprietary, with task-complexity classification ahead of selection | Constant cognitive load. Users face five intelligence tiers and no basis to choose, especially for knowledge work rather than coding. One member described “stacking up model debt.” | Strong |
| Unit economics frameworks for AI | Unit cost matters more than in cloud precisely because value is harder to establish. | Strong |
| Token optimization playbooks | Efficiency levers exist but are scattered and undocumented. | Moderate |
| Accounting and finance guidance: is token spend COGS, can it be capitalized, how do you account for agentic labor, what if an agent creates an agent | Reporting processes are greenfield and reach back into procurement. Raised as a top ask by a finance leader at a very large institution. | Moderate |
| Allocation and entitlement policy guidance: usage tiering by seniority, by job family, by workload complexity | Companies are doing it implicitly and denying it. Blanket caps were named as the wrong answer. | Moderate |
| A business case question set for open-weights-on-owned-hardware proposals | Buyers have never run AI workloads in their own data centers and will not know what to ask when the proposal lands. Suggested lines of inquiry: opportunity cost of building, neocloud as a faster alternative, short contracts instead of large commitments. | Single |
| A multi-year evolution roadmap: what this looks like today, in three years, in five, in ten | Nothing is commoditized yet, and the math changes as it commoditizes. | Single |
| Honest guidance on where this will and will not work, given that most organizations lack the integrated structure the framework assumes | Explicitly framed as avoiding the FinOps trap of prescribing something organizations cannot execute. | Single |
This is how you should manage your AI cost. This is our prescription to do that. Obviously it will get refined and matured over conversations, but that guiding principle needs to be laid down.
I definitely think frameworks would be quite helpful. There’s just so much choice, and how do you even get started is part of the journey for most companies.
Tier 5. Education, training, certification
This tier had the widest range of requested altitudes, from “what is a token” up to certifying an individual as qualified to use frontier models.
-
Baseline literacy
The floor: what a token is, what a model is, what an agent is. Multiple members reported watching enterprise audiences unable to answer these. One noted the scale problem plainly: cloud meant tens of thousands of engineers needed to learn something; AI means hundreds of thousands of employees do. Another cited sixty thousand no-code users being onboarded to a frontier model at a single institution. Strong -
A level-one course and certification
The most-requested single education artifact. The pain has two halves. Engineers: “I don’t even know what I don’t know, it would be great to have a course that walks me through all of the concepts.” Leaders: they are bluffing in meetings. One member’s finance team specifically valued the equivalent FinOps course precisely because it served non-technical staff. Delivery model discussed as thin hosted content plus curated external links, with the certification carrying the weight. Strong -
Two-audience education: corporate and end user
Requested as an explicit pair. Corporate-level guidance on where to invest and how to govern, plus end-user guidance on how much to use and for what. Moderate -
Curated external reading list
The constraint is not availability but knowing which of thousands of sources is worth reading. “Our course would be good at navigating you to good, relevant material out there.” Single -
Deeper technical tiers
For the segment that needs depth the general course cannot carry. The governing principle proposed: teach the concept, not this week’s release. Release-level detail belongs in articles, not curriculum. Moderate -
A book, with the course derived from it
Chapter by chapter, on the precedent that the original FinOps certification was structured on a book outline. Single -
Certification as a gate on model access
The most provocative ask in the set: a badge that certifies an individual is qualified to use a given tier of model. Motivated by a real waste pattern, unqualified users routing trivial questions to the most expensive models, and by the absence of any prompt standards. The requester framed the incentive candidly: a badge that says you are good enough to use the best models is something people would pursue for career reasons. Independently, banks asked whether model access could be gated behind a certification layer for non-technical users. Moderate -
Long-tail business user training
Input versus output token cost, document format cost, image versus text cost. Practical hygiene at the consumption layer. Unresolved whether this belongs to tokenomics or FinOps. Moderate -
An adoption pathway for newcomers
A meaningful share of inbound is still at a 101 level, especially outside North America. The Foundation risks talking past them entirely. Single
There’s just a mass amount of education for people. There’s a corporation side of that, but there’s a user education side of that as well.
If the badge is saying you know this stuff, that’s the important thing, because that sets the tone of what it means to be certified.
Tier 6. Information radiation
Distinct from standards work, and explicitly so. Members want fast, current, opinionated signal on a space that changes weekly, delivered while the slow standards work proceeds. One member named the gap precisely: plenty of content mentions tokenomics as a topic, but nothing treats it as the topic.
| Channel | What was said | Signal |
|---|---|---|
| Newsletter | Requested directly, with a stated preference for focused and concise. Consistently the least contested channel. | Strong |
| Podcast | Wanted by leadership-facing members, who named podcasts as where their peers and executives consume this. One member volunteered to appear. Another explicitly ranked podcasts lower than a newsletter for their own use, so the demand is real but segmented. | Moderate |
| Presence on X | Named twice as where the actual technical discourse happens, and where the deep practitioners write. One observation: ask a FinOps person where they get information and they say a professional network; ask a tokenomics person and they say X. | Moderate |
| Professional network reach | Still the widest reach and where one member gets nearly all their information, but characterized elsewhere as surface-level and increasingly flooded with generated content. | Moderate |
| White papers | Named as materially useful precedent from the FinOps Foundation, and requested for replication here. | Strong |
| Reactive insight articles on new techniques and announcements | The release-level counterpart to the concept-level curriculum. | Moderate |
| Academic literature and papers | There is nothing to build on. Co-authored papers and cross-referencing open-source frameworks were both proposed. | Single |
A hard constraint on all of it
Tier 7. Community, convening, peer access
- Peer networking with comparable organizations. The most consistently stated reason for joining. Members want to know what companies of their scale are actually doing, not what vendors say. Strong
- Anonymized story sharing, wins and losses. Including provider-to-provider. The stated alternative is waiting for analyst research or generalizing from one unrepresentative deal. One member committed to contributing anonymously. Moderate
- Working group participation paths for people who do not report to the member sponsor. AI capability sits outside the sponsor’s org in matrixed companies, and the people who need to be in the room cannot be assigned by the board member. Single
- Member-to-member introductions and mentoring, including access to advanced practitioners for research purposes. Moderate
- Events and meetups outside North America, specifically Asia Pacific and Southeast Asia, where inbound demand was reported as high and local presence as absent. Single
- Cross-foundation alignment work, including a working group that brings in the adjacent agentic AI effort. The concern raised was pointed: one AI foundation setting a standard that contradicts another’s, and standards from safety and security work that are not reusable here. Framed as a strategic risk, not a nicety. Single
- Speaking and thought leadership slots for members. Several members want to present their point of view; at least one wants credibility as a contributor, not revenue, as the entire return on membership. Moderate
- University and academic engagement, including associate membership routes and research collaboration. Moderate
We are just feeling the elephant right now. Learning and networking from other bigger companies like ours, that’s number one. What are they doing?
She very much wants us to invest in this in a way that we are seen as a contributing member and helping define what those best practices and frameworks are. It wouldn’t be about revenue. It would be, are we contributing enough, and of enough quality, that we are seen as having a voice here.
Tier 8. Open source and hosted projects
- Host external open source projects in the way a technology foundation normally does, cited as the model that pulled an earlier fragmented orchestration market into coherence. Cache and KV cache tooling was named as the concrete example of a project worth standardizing on. Notably, the same member declined to nominate specific projects yet: “I absolutely want to bring a couple of the projects. I just don’t know. I really want to see where they go.” Single
- Open source calculators and primitives, for example token complexity estimation, as starting places for others to build on rather than products. Moderate
- Publish open standards and open data to raise market maturity, with at least one member willing to contribute their own data in open form. Single
- A task complexity classification model for sizing agentic workloads ahead of model selection, already donated and close to publication. Moderate
- Conformance or attestation that an organization is practicing tokenomics, as a downstream consequence of metering standards and taxonomy. Proposed rather than requested. Single
Explicit non-goals
What we said we would not do
- Build products. No model routing, no intelligence layer, no optimization tooling.
- Publish model benchmark scores. Supply the metrics and the framework so members benchmark their own workloads.
- Teach general AI concepts. The precedent is declining to teach cloud fundamentals. This is under real pressure from members who report their audiences cannot answer “what is a token.”
- Recommend vendors. Neutrality is the asset.
- Become a pointer service. Pushing too far up the abstraction stack collapses the discipline back into FinOps.
What is contested about this roadmap
- Scope breadth. One member’s advice was to pick a single lane and demonstrate visible value inside twelve months, on the grounds that a foundation covering everything is “just another comet in the sky.” Another rejected picking lanes at all, arguing production, consumption, and value are one loop. Both were serious, and they cannot both be followed.
- Standards speed. A platform vendor’s objection stands unrebutted: a six-month standard may be obsolete on arrival, and the opportunity cost of member time is real. Their own counterweight is the risk of a standard emerging that they did not shape.
- Value first or last. Some members ranked value as the highest-differentiation work. Others said ROI modeling is too sophisticated for a market that still cannot attribute basic cost, and that value will not be solved in months. One drew a hard line and put value outside the token question entirely.
- Whether education is in scope. Baseline literacy is heavily requested and explicitly declared a non-goal. That contradiction needs resolving, most likely by deciding whether the FinOps Foundation or the Tokeneconomics Foundation owns it.
- Prescription versus documentation. Members asked for a prescriptive framework. Members also said every company attributes value differently and there is no one-size-fits-all. The likely landing spot is documenting the valid approaches rather than mandating one, but that is not what was asked for.
The near-term consensus
Stripping out everything contested, five items had broad support, low dependency on other work, and were described by at least one member as achievable this year.
- Publish a definition of tokenomics. Mineable directly from these conversations. Nothing else is blocked on new research.
- Publish a definition of a token and the token types. Named first by more members than any other single item.
- Publish the taxonomy and the component cost picture. The vocabulary everything downstream depends on.
- Ship a cost measurement recipe. Cost only. Value deliberately deferred.
- Start the newsletter. Lowest cost, least contested, and it fills a gap members named explicitly while the slow work proceeds.
One caution