CostGrid Agentic FinOps & AI governance
LLM inference cost governance & unit economics

Per-token prices are collapsing.

Your AI bill is going up anyway.

Unit price is the wrong thing to watch. Spend is price multiplied by volume, and volume usually wins. CostGrid is a dashboard and forecast model for the real cost of running AI in production — the token line, the governance line, and the routing decision that sits between them.

OPEN THE LIVE DASHBOARD → DOWNLOAD THE MODEL · XLSX
Enterprise GenAI spend
2026 $69.1B
2030 $207.3B

Prices fall roughly 10× a year. Spend triples. Falling prices don’t shrink the bill — they grow it.

21.8%
of head-to-head pairs, the cheaper-listed model costs more — by up to 28×
1 in 5
firms has mature agent governance — Deloitte
97%
of AI security incidents lacked access controls — IBM
Aug 2026
EU AI Act fines begin
01 — Why now

Tokens are the visible line, not the whole bill.

Cost taxonomy

The token line is separated from governance, deployment, observability and change-management cost, then rolled into a levelized cost per valid inference.

Routing economics

Cheaper models are not automatically cheaper. Open weights cost far less per token but carry misrouting risk that grows convexly as you push high-stakes work onto them.

Price forecast

Tiers decay toward a hardware floor at different rates. Economy and open-weight halve every ~1.1 years; frontier reasoning resists the curve almost entirely.

LCOAI = (amortized CapEx + total OpEx) ÷ inferences that cleared policy
02 — The dial

Stop forecasting vendor prices. Forecast your routing.

Substitution share is the percentage of traffic routed from frontier models down to the open-weight floor. It is the one input an organization fully controls. Move it.

Substitution share 15%
0% 99%
Blended cost / M tokens
Monthly token spend
vs. routing 15% today

Blended cost curve · $/Mtok
$11.00 $6.75 $2.50 today 15% optimum 83% s = 0% s = 99%

Cost falls as routing shifts to the open floor, then the s8 misrouting-risk penalty bends it back up past ~90%. The optimum is where the marginal token saving equals the marginal risk penalty — not wherever the cheapest vendor sits.

Commodity work
Ticket classification
frontier → open floor
Bounded risk
Contract summarization
open pro, mid-tier fallback
Accountable core
M&A diligence analysis
stays frontier — trust is scarce
03 — Where value accrues

Value never accrues where prices converge.

Trust & accountability

Governance, security, sovereignty, the accountable boundary. Scarce, sticky, and where margin and exit value accrue.

margin ↑
Orchestration & routing

Where substitution share lives: model-tier-to-task, evals, output caps, fallback. The lever you control.

the dial
Capability — models & tokens

Deflating toward the cost of electricity. Commoditizing, with no durable margin.

price ↓

CostGrid operates the top two layers.

04 — Forecast

2026 → 2030, bifurcation scenario

Six modules driven entirely from an assumptions sheet: price decay, routing, volume, spend, seat pricing, EBITDA impact. Change an input and it propagates through to the exit bridge.

Metric 2026 2030
Frontier blended $/Mtok $10.00 $7.20
Open-weight floor $/Mtok $0.35 $0.10
Frontier / open gap 28.6× ~70×
Risk-adjusted optimal share 83% 90%
Enterprise GenAI spend $69.1B $207.3B

Illustrative organization, mid-2026 snapshot. Prices and model versions move monthly — re-verify against official pricing pages before external use.

05 — Inside the dashboard

One self-contained file. No build step.

01Agent cost explorer — 14 agents, cost-of-pass and useful-token ratio per row
02Substitution-share control with live blended-cost curve
03Live enforcement feed — monitor / warn / block per agent
04Pricing reference with project-verified anchors
05Exportable executive summary and EBITDA & exit bridge
OPEN THE DASHBOARD →
[ dashboard screenshot ]
drop a 1600×1000 capture here
06 — Method

Three choices worth stating plainly, since they drive most of the output.

k · f · sN, N = 8

The risk penalty on misrouted work. A high exponent encodes the assumption that routing errors are cheap at the margin and very expensive in the tail.

logistic volume curve

Tokens per interaction follow a curve with a ceiling, not a growth rate. The shift to multi-step agentic orchestration is a pattern change, not a compounding trend.

hardware floor > 0

Prices decay toward electricity-and-GPU benchmarks, not toward zero. No tier declines past what inference physically costs to serve.

Grounded in NIST AI RMF / EY — The Total Cost of Agents / Stanford HAI AI Index 2026 / Citadel Securities / 12 papers →

Tokens are cheap. Trust is not. Govern the substitution.

OPEN THE LIVE DASHBOARD → VIEW THE REPOSITORY
CostGrid — built by Tomas Buica MIT licensed · figures illustrative, not a real organization