AI TOKENOMICS · ASSESSMENT

How we run a tokenomics assessment.

A tokenomics assessment shows you where your AI spend goes, where it is heading, and where you can bring it down — measured on your own traffic, inside your own cloud. This page walks through what we examine, how we measure it, and what you keep at the end.

An AI bill arrives as a single number. It doesn't tell you which feature, customer or job spent it — and that mix keeps shifting month to month as work moves onto heavier models and agent loops multiply, often while the total stays flat enough to go unnoticed.

We measure your spend inside your own environment, track where it is heading over time, and put a dollar figure on each fix. In short: we make the spend visible, score it, project it, price the fixes, and prove the easy ones before we leave.

00WHAT YOU GET

The four deliverables.

Everything comes out of one measured pass over your own data, so the breakdown, the forecast and the plan all trace back to the same evidence. You keep all four when we leave.

01

A living spend breakdown

Your AI bill attributed to features and jobs as far as the logs allow — with whatever cannot be attributed shown on its own line, tied back to the invoice finance already reconciles, and read across time rather than as a single-week snapshot.

02

A scorecard

Eleven signals across four lenses, each given a rating and a trend, and each gap turned into a dollar figure with the evidence attached.

03

A trend map and a do-nothing forecast

How your spend is moving, one axis at a time, with a two-to-four-quarter projection of the bill if nobody touches anything.

04

A plan to bring it down

Where there are easy, low-risk wins, we take them during the run and prove them on your own traffic. Beyond that, a prioritised list of the larger changes worth considering — each one sized so you can decide what is worth doing. All against a cost-per-outcome target you keep.

If your spend is already lean and the curve is flat, we say exactly that — a one-page all-clear — and stand down. The verdict is whatever the data returns, including "there's nothing here to do." We do not recommend a second phase the data does not support.

·HOW WE WORK

Four principles behind every number.

These are what make the numbers trustworthy. They hold from first access to final readout — each one a rule you can hold us to.

01

Test on your own traffic

Every model and every fix is decided by re-running your own traffic against the alternatives and watching what happens — tokens, dollars, latency, whether the answer still holds. A public benchmark does not show where your workloads land.

02

Every number shows its source

Nothing is shown untagged. Each figure is marked by how solid it is — measured, logged, modelled, or assumed. A combined figure inherits the weakest tag of its inputs, and anything modelled ships as a range.

03

Price the task

The unit that matters is one finished piece of work — which may be many calls, tool hops and retries stitched together. Per-token totals hide the jobs that cost the most.

04

Nothing leaves your cloud

Reading logs, redaction and testing all run inside your own account and region; tests hit your endpoints on your quota. No prompt, response or invoice is copied to us. This is in the contract.

01SCOPE & ACCESS

What's in scope, and the access we need.

The results are only as honest as the data underneath them, so we settle two things before starting: where the assessment begins and ends, and exactly what read-only access you give us.

Where the line sits
We coverEvery model API call — direct, cloud-hosted, or through a gateway — plus agent workloads, embeddings and reranking.
We don'tSelf-hosted GPU economics, training and fine-tuning runs, and quality work unrelated to cost. We name these and move on.
We route outFine-tuned or heavily customized models — detected, flagged, and handed to specialists rather than tested blind.
Access we need up front
BillingRead-only on provider consoles and usage exports — six months if it exists, so the trend read is real; 30–90 days is the floor.
LogsRequest logs where they're kept, plus a map of which key belongs to which service. If logging is off, that's finding number one — we run a ~1-week capture first.
ArchitectureA diagram per AI-calling service and the name of each feature's owner, for interviews and sign-off.
Data termsA signed note: processing stays in your environment, customer data redacted at ingestion, nothing leaves. Principle four, on paper.
The run, end to end
01Accessscope · owners · data terms
02Measure & readrebuild the past + meter the present
03Four lensesrate 0–3 · trend
04Trend read8 axes · forecast
05Price the fixes3 curves · ranges
06Ship & verifyprove it in your data
07Readout90 min · one room
02MEASURE & READ

How we measure the spend.

We do two things at once: rebuild the last two quarters of spend from your invoice inward, and start measuring live traffic going forward. Both are organised by job and laid out over time — so that the change in spend is visible, not only the total.

01

Group by job

A workload is the unit every later decision hangs on — a cluster of requests doing the same thing, like one product feature or one agent. When no clean list exists, we rebuild one: the label the app emits, else the caller's identity, else clustering prompts by their fixed structure. Anything that still cannot be attributed sits in a named bucket.

02

Rebuild the spend, month by month

The invoice sets the total finance trusts. Then we split each job into input / output / cached / reasoning — because that mix is what identifies the fix. Done month by month, it shows early movement: work climbing to heavier models, ratios drifting, agent loops multiplying, reasoning tokens bloating.

03

Meter live traffic

A gateway and a trace layer — whichever you already run — sit in front of traffic. Every call is stamped with who issued it, where, and the two fields that carry the weight: the task it belongs to (which fuses many calls into one unit of work) and the turn number (which exposes context growth).

Close-out check: the rebuild has to meet the invoice within 5%, or the difference is named — untracked keys, marketplace lines, rounding. We call the read done when tagging covers >90% of traffic, the draft breakdown reconciles within 5%, every job has a named owner, and at least two months of history are in hand to read the trend.

03FOUR LENSES · ELEVEN SIGNALS

What we score, and how.

With the spend finally visible, we check it against four groups of questions — four lenses, eleven signals in all. Each signal gets two readings: a rating on a four-rung ladder (Blind → Visible → Under control → Improving on its own), and a trend showing whether the gap is widening, holding, or closing on its own. The priced gap and its direction matter more than the rating.

Lens A

Can you see it?

Can you see the spend, and see it move?
A1

Traceable spend

Can every dollar on the invoice be traced to a job, a team, and what it bought? Every later figure depends on it.

How we read itReconcile tagged spend to the invoice; attribute the untagged tail by who owns the keys and by log volume. Many AI-native teams meter nothing below the invoice line, so coverage is the first gap to close.

0Blindinvoice is the only number
1Visibleprovider / key totals
2Under controlper-job, most traffic
3Improving on its ownper-outcome, self-maintaining

FixesUnlocks every other lens, plus per-customer margin & pricing.

A2

Cost per outcome

Does the setup know the price of one delivered unit — a resolved ticket, a completed task — and is that price tracked and used?

How we read itJoin usage to outcome events; where the join is missing, stand up a feature-level proxy and mark it modelled. Set unit cost against the outcome's value to read headroom or alarm.

0Blindnever computed
1Visiblecomputed once, now stale
2Under controlfeature-level proxy, tracked
3Improving on its owntrue unit cost, drives pricing

FixesOutcome-priced billing, invest / retire calls, a margin-aware roadmap.

A3

How fast you catch a spike

How long a cost spike lives before anyone notices — a prompt edit that triples context, a retry storm, an agent that forgets to stop.

How we read itWalk the last quarter for step-changes and, for each, separate when it began from when it was caught. If the record is silent, we stage a controlled spike and see what fires.

0Blindcaught on the invoice
1Visiblecaught in weekly review
2Under controlsame-day automated alerts
3Improving on its ownalerts + cost checks in CI

FixesCost alerts, per-key ceilings, cost-diff checks at deploy.

Lens B

Right-sized models?

Is each job on the cheapest model that still holds up?
B1

Right-sizing

Whether each job runs on the cheapest model that still clears its quality bar — proven by replaying its own traffic. Top-tier to small pricing spans 10–25×; most extraction and classification never needed the top rung.

How we read itTake a representative slice of the job's own prompts and run it down an ordered set of models, cheapest first. Fire each a fixed number of times and grade every result against your own production output. Stop at the first one that holds. Anything that clears nothing goes to human review with its evidence.

0Blindone flagship for everything
1Visiblead-hoc tiering, no evidence
2Under controltest-mapped, static
3Improving on its ownre-mapped whenever prices move

FixesStatic remapping — the highest-yield, lowest-effort move — then smaller custom models for narrow high-volume work.

B2

Smart routing

Whether the setup routes by difficulty — cheap model first, escalate only what fails — or sends everything to one tier regardless of how hard the request is.

How we read itFrom the right-sizing tests, isolate jobs where a cheap model passes most of the time but not always — the cascade candidates. Model the blended cost of cheap-first-plus-escalate against the flat-flagship baseline, and price the difference.

0Blindno routing, one tier
1Visiblemanual split by endpoint
2Under controlconfidence cascade in place
3Improving on its ownverifier-gated, self-repricing

FixesConfidence cascades, verifier-gated escalation, provider-fallback routing.

Lens C

What's in each call?

What rides inside each call, and is it bloated?
C1

Context reuse

The shape and reuse of everything sent into the model. A stable prompt prefix can be re-read by the provider at up to a ~90% discount, so what matters is how much of each call is reused rather than freshly billed. Agent workloads run 100:1–300:1 input-to-output because the full context is resent every turn — which makes reuse the main lever.

How we read itRead cached-vs-fresh input straight from provider usage; analyse sampled templates for the reuse ceiling and the cache-breakers — volatile data spliced ahead of stable scaffold, tool lists that reorder, per-session IDs in the system prompt. Price the distance to the ceiling at the cached rate.

0Blindno reuse, profile unknown
1Visibleaccidental hits, known in aggregate
2Under control>60% reuse, structured prefixes
3Improving on its own>85% + a semantic layer where it pays

FixesStatic-prefix restructuring, cache-TTL tuning, an app-level semantic cache.

C2

Output control

Control on the expensive side. Generated tokens price at 5–6× input, and reasoning tokens bill as output — often at a high default nobody set on purpose. The most common overspend we find.

How we read itAudit request settings at every call site in code, joined to trace data. Compare what the model produces to what the downstream code consumes — three parsed fields don't need a paragraph wrapped around them. Inspect reasoning settings explicitly; unset or defaulted-high is where it leaks.

0Blindall defaults, prose everywhere
1Visiblecaps on a few endpoints
2Under controlcaps + structure + budgets standard
3Improving on its ownenforced in review & CI

FixesOutput ceilings, structured outputs, per-class reasoning budgets, terser prompts.

C3

Agent-loop cost

Tokens consumed per completed task inside agent loops, where each turn re-sends the whole conversation and tool output so far. It's the fastest-growing line in most setups: a ten-turn task can cost ~50× a single call, and trimming stale context alone recovered ~84% in a published 100-turn evaluation.

How we read itRebuild each loop from grouped traces and measure what fills the window turn by turn — instructions, tool schemas, history, results. Surface the usual offenders: full schemas resent for a two-tool task, multi-KB dumps never read again, no compaction, no stop condition.

0Blindno visibility into loops
1Visiblecall counts known, composition not
2Under controltrimming + compaction in place
3Improving on its owntokens-per-task budgeted & owned

FixesTool-result trimming, scheduled compaction, sub-agent isolation, loop budgets.

Lens D

Discounts & guardrails

Are you paying list price, and are there hard limits?
D1

Discounts claimed

Whether you pay list price when you need not. Discounts already on the table — batch and flex tiers (a slower lane, roughly half price, for work no user is waiting on), caching, volume commitments — frequently go unclaimed. Prices fall ~10×/year at a fixed capability, so a routing choice two quarters old was made against a price sheet that no longer exists.

How we read itSize the eligible discount per class from the inventory, classify each job by how long it can wait, and price async volume stuck on the realtime tier at the flat ~50% recovery. Then find who owns the re-pricing cadence — usually no one.

0Blindlist price on everything
1Visibleone discount class captured
2Under controlmost classes, annual review
3Improving on its ownall classes, quarterly re-pricing

FixesBatch / flex migration, commitment negotiation, a refresh cadence tied to releases.

D2

Wasted spend

Money on calls that returned nothing usable — hard failures, re-parse retries, duplicate submissions, double-completed timeouts.

How we read itRead status codes, retry headers and near-identical prompts in tight windows from gateway logs, then attribute cause — re-parse loops point at missing structured output, 429 bursts at absent backoff, duplicates at missing idempotency. A well-run setup sits in the low single digits.

0Blindnever measured
1Visible>5%, causes unknown
2Under control2–5%, causes known
3Improving on its own<2%, watched continuously

FixesStructured outputs, idempotency keys, jittered backoff, circuit breakers.

D3

Spend limits

Whether spend has hard stops — per-key and per-team ceilings, rate limits, a kill-switch that fires — and whether anyone owns the page when one trips.

How we read itAudit gateway and provider config, then prove it — drive a test key into its limit and confirm the block, rather than trusting the setting. Ask who got paged the last time a ceiling tripped, and what happened next.

0Blindno limits anywhere
1Visibleprovider cap only
2Under controlenforced per-key budgets
3Improving on its ownenforced + tested + paged owner

FixesEnforced per-key ceilings, dev / prod key separation, per-job spend targets.

04THE TREND READ

Where your spend is heading.

This is the part a one-time snapshot can't give you, and the reason the whole assessment runs over time. Your spend is never still — model line-ups change, agent features spread, prompts grow, providers cut prices. A figure that was acceptable two quarters ago can be an overspend today. We track each of these movements and project it forward.

Climb to heavier models

Spend drifting up the model stack. Usually one team adopts a new reasoning model and it seeps into jobs that never needed it.

Agent share

The slice of the bill sitting inside agent loops. A flat total can hide an agent share quietly doubling under it.

Context growth

The per-turn growth curve over time. Steepening month over month means history-replay cost is compounding.

Reasoning share

Thinking tokens as a rising fraction of output — the quiet drift as more work lands on reasoning-capable models.

Reply-length creep

Typical reply length per job over time, held against what the consuming code reads.

Waste trend

Waste as a ratio tracked over time. Retry storms and duplicates tend to grow with traffic until they are noticed.

Unit-cost curve

Cost per outcome against volume growth. Rising unit cost on rising volume is the signal finance watches for.

Price-sheet risk

Provider repricing and new releases you haven't reacted to — the outside axis moving whether or not anyone looks.

The do-nothing curve: from the measured motion we extend the bill two-to-four quarters out on the assumption that nobody intervenes — always modelled with a range, never a single line. A flat total often hides a changing mix underneath, and moving every job to the newest model often raises the bill. The per-job distribution is where the value is.

05PRICING THE FIXES

Putting a dollar figure on each fix.

Every gap and trend signal is turned into money the same way, each line carrying its evidence tag. We price each fix twice — against what you spend today, and against the do-nothing forecast you're already heading toward — because that second, growing number is what shows the real cost of waiting.

Do-nothing

Where it lands untouched

Where the trend read puts the bill if nobody acts. Every saving below is measured against this line as well as against today.

modelled · range
Right model

Test-proven

Every job on the most economical model it cleared in the tests. This curve is measured.

measured
Model + payload + discounts

Everything the profile unlocks

Fixes applied in order — route, then cache, then batch, then trim — because they stack, they don't add.

stacked

Rule: fixes stack, they don't add. Route a job to a model 10× cheaper and you've shrunk the base that caching then discounts — so they are applied in sequence, never summed, or the total is overstated. Every recommendation also carries a build estimate; anything paying back beyond two quarters is roadmap, not a quick win, and it's labelled as such.

06SHIP & VERIFY

What we ship during the run.

Most reviews end with a list of recommendations. Where we find easy, low-risk wins during the run, we take them — implement them and verify the saving on your own traffic — so the readout opens with proven results. If nothing is safe to ship, we say so.

If a rebuild is worth it — the order it goes in
1Evaluations first, so routing is safe.
2Routing & cascades, built on top of the evals.
3Agent context engineering and semantic caching.
4Smaller custom models last, for stable high-volume tasks.
What you keep running it
ProofBefore-and-after on the same traffic, with a holdout where volume allows. A saving is claimed only once it shows up in the gateway data.
The targetA cost-per-outcome model in your hands: today, after the quick wins, after the rebuild — plotted against volume growth and against the do-nothing curve.
Operating modelDashboards, a budget policy with named owners, a monthly review. Two standing rules: every provider price change triggers a re-routing pass, and the trend axes are watched so slippage is caught in-month, not on the invoice.
07THE READOUT

The final readout.

It all lands in a single ninety-minute session, with engineering and finance in the same room — so both sides work from one set of numbers. The running order:

0–15
The readWhat you spend, where it goes, and how the shape has moved across the window.
15–35
The scorecardWhere you stand across the four lenses — every rating and trend backed by the query behind it.
35–50
The trend & do-nothing forecastWhere the spend is heading, and what the bill becomes if nobody acts.
50–65
The priced fixesWhat each gap is worth — evidence-tagged, three curves, a floor and a likely number.
65–75
What already shippedThe wins landed during the run, proven in your own gateway data.
75–85
Rebuild & targetThe sequenced plan and the curve you could be on.
85–90
Phase two, if there is oneOutcome-linked commercials, anchored on the floor.

And if the read comes back lean and the curve comes back flat, this meeting is ten minutes and an all-clear. We would rather report that than recommend a rebuild you do not need.