Unit price is the wrong thing to watch. Spend is price multiplied by volume, and volume usually wins. CostGrid is a dashboard and forecast model for the real cost of running AI in production — the token line, the governance line, and the routing decision that sits between them.
Prices fall roughly 10× a year. Spend triples. Falling prices don’t shrink the bill — they grow it.
The token line is separated from governance, deployment, observability and change-management cost, then rolled into a levelized cost per valid inference.
Cheaper models are not automatically cheaper. Open weights cost far less per token but carry misrouting risk that grows convexly as you push high-stakes work onto them.
Tiers decay toward a hardware floor at different rates. Economy and open-weight halve every ~1.1 years; frontier reasoning resists the curve almost entirely.
Substitution share is the percentage of traffic routed from frontier models down to the open-weight floor. It is the one input an organization fully controls. Move it.
Cost falls as routing shifts to the open floor, then the s8 misrouting-risk penalty bends it back up past ~90%. The optimum is where the marginal token saving equals the marginal risk penalty — not wherever the cheapest vendor sits.
Governance, security, sovereignty, the accountable boundary. Scarce, sticky, and where margin and exit value accrue.
Where substitution share lives: model-tier-to-task, evals, output caps, fallback. The lever you control.
Deflating toward the cost of electricity. Commoditizing, with no durable margin.
CostGrid operates the top two layers.
Six modules driven entirely from an assumptions sheet: price decay, routing, volume, spend, seat pricing, EBITDA impact. Change an input and it propagates through to the exit bridge.
| Metric | 2026 | 2030 |
|---|---|---|
| Frontier blended $/Mtok | $10.00 | $7.20 |
| Open-weight floor $/Mtok | $0.35 | $0.10 |
| Frontier / open gap | 28.6× | ~70× |
| Risk-adjusted optimal share | 83% | 90% |
| Enterprise GenAI spend | $69.1B | $207.3B |
Illustrative organization, mid-2026 snapshot. Prices and model versions move monthly — re-verify against official pricing pages before external use.
The risk penalty on misrouted work. A high exponent encodes the assumption that routing errors are cheap at the margin and very expensive in the tail.
Tokens per interaction follow a curve with a ceiling, not a growth rate. The shift to multi-step agentic orchestration is a pattern change, not a compounding trend.
Prices decay toward electricity-and-GPU benchmarks, not toward zero. No tier declines past what inference physically costs to serve.
Tokens are cheap. Trust is not. Govern the substitution.