Featured Project
The five-layer tokenomics stack
Open a layer to see what it controls
L5
Routing and governance
bounds fan-out
The control plane in front of everything else: which model serves each request, budgets and quotas, and the guardrails on agent depth and retries. The only layer that can decide a request needs no model at all.
Lever: complexity routing, budgets, agent caps, circuit breakers.
L4
Model and quantization
right-sizes what runs
Which model runs, at what numeric precision, with what adapted weights. Running a frontier-scale model for a trivial task is the classic waste.
Lever: right-size the model, quantize, adapt.
L3
Inference stack
suppresses repeated work
The serving software: engine, KV cache management, batching, prefill and decode split. Usually the single largest source of optimization, and the first layer where consumption is suppressed rather than merely priced.
Lever: right engine, prefix and KV cache, disaggregation.
L2
Capacity and Energy
sets effective cost
How much silicon you hold, where it sits, and how much of it does useful work. Idle capacity raises the effective cost of the work that does run.
Lever: batch the troughs, autoscale, place by region.
L1
Silicon
sets the floor
The accelerator and its generation, with the memory, interconnect, and power it ships with. It converts energy into tokens and sets the floor on cost per token for everything above it.
Lever: match silicon class to workload, adopt newer generations.
Featured Project
Big-T Notation
Big-O for AI. A shared language for how token consumption grows as usage scales.
T(1)
Constant. The model is not called per request — a cache hit, a static lookup.
cache and precompute
T(log n)
Sublinear. Deterministic code shrinks the input before the model sees it.
filter before inference
T(n)
Linear. One model call per request — the healthy default.
trim per-call overhead
T(n·k)
Multiplicative. k model calls per request, and k is usually invisible.
compose tool pipelines
T(n·k·a)
Agent-multiplicative. An orchestrator spawns sub-agents that spawn tool calls.
bound depth, add budgets
T(∞)
Unbounded. Loops with no termination condition.
hard termination. always.
T(n · k · a) n = requests or input size · k = model calls per request · a = agent depth