AI TOKENOMICS · ASSESSMENT

How we run a tokenomics assessment.

A tokenomics assessment shows you where your AI spend actually goes, where it's heading, and where you can bring it down — measured on your own traffic, inside your own cloud. This page walks through what we examine, how we measure it, and what you keep at the end.

An AI bill arrives as a single number. It doesn't tell you which feature, customer or job spent it — and that mix keeps shifting month to month as work moves onto heavier models and agent loops multiply, often while the total stays flat enough to go unnoticed.

A normal audit reports that number and stops. We go further: we measure your spend inside your own environment, track where it's heading over time, and put a dollar figure on each thing worth fixing. Everything below is how that runs — and what you get. In short: we make the spend visible, score it, show where it's going, price the fixes, and — where there are easy wins — prove a few before we leave.

00WHAT YOU GET

The four deliverables.

Everything comes out of one measured pass over your own data, so the breakdown, the forecast and the plan all trace back to the same evidence — not three separate exercises that disagree. You keep all four when we leave.

01

A living spend breakdown

Your AI bill attributed to features and jobs as far as the logs allow — with whatever can't be cleanly attributed shown on its own line, not forced into a bucket — tied back to the invoice finance already reconciles, and read across time rather than as a single-week snapshot.

02

A scorecard

Eleven signals across four lenses, each given a rating and a trend, and each gap turned into a dollar figure with the evidence attached.

03

A trend map and a do-nothing forecast

How your spend is moving, one axis at a time, with a two-to-four-quarter projection of the bill if nobody touches anything.

04

A plan to bring it down

Where there are easy, low-risk wins, we take them during the run and prove them on your own traffic. Beyond that, a prioritised list of the larger changes worth considering — each one sized so you can decide what's worth doing, rather than a promised re-architecture. All against a cost-per-outcome target you keep.

If your spend is already lean and the curve is flat, we say exactly that — a one-page all-clear — and stand down. The verdict is whatever the data returns, including "there's nothing here to do." We don't manufacture a problem to justify a second phase.

·HOW WE WORK

Four principles behind every number.

These are what make the numbers trustworthy. They hold from first access to final readout — each one a rule you can hold us to.

01

Test on your traffic, not a leaderboard

Every model and every fix is decided by re-running your own traffic against the alternatives and watching what happens — tokens, dollars, latency, whether the answer still holds. A public benchmark tells you nothing about where your workloads actually land.

02

Every number shows its source

Nothing is shown untagged. Each figure is marked by how solid it is — measured, logged, modelled, or assumed. A combined figure inherits the weakest tag of its inputs, and anything modelled ships as a range, not a single number.

03

Price the task, not the token

The unit that matters is one finished piece of work — which may be many calls, tool hops and retries stitched together. Per-token totals flatter the bill and bury the jobs that actually cost.

04

Nothing leaves your cloud

Reading logs, redaction and testing all run inside your own account and region; tests hit your endpoints on your quota. No prompt, response or invoice is ever copied to us — it's in the contract, not left to trust.

01SCOPE & ACCESS

What's in scope, and the access we need.

The results are only as honest as the data underneath them, so we settle two things before starting: where the assessment begins and ends, and exactly what read-only access you give us.

Where the line sits
We coverEvery model API call — direct, cloud-hosted, or through a gateway — plus agent workloads, embeddings and reranking.
We don'tSelf-hosted GPU economics, training and fine-tuning runs, and quality work that isn't about cost. We name these and move on.
We route outFine-tuned or heavily customized models — detected, flagged, and handed to specialists rather than tested blind.
Access we need up front
BillingRead-only on provider consoles and usage exports — six months if it exists, so the trend read is real; 30–90 days is the floor.
LogsRequest logs where they're kept, plus a map of which key belongs to which service. If logging is off, that's finding number one — we run a ~1-week capture first.
ArchitectureA diagram per AI-calling service and the name of each feature's owner, for interviews and sign-off.
Data termsA signed note: processing stays in your environment, customer data redacted at ingestion, nothing leaves. Principle four, on paper.
The run, end to end
01Accessscope · owners · data terms
02Measure & readrebuild the past + meter the present
03Four lensesrate 0–3 · trend
04Trend read8 axes · forecast
05Price the fixes3 curves · ranges
06Ship & verifyprove it in your data
07Readout90 min · one room
02MEASURE & READ

How we measure the spend.

We do two things at once: rebuild the last two quarters of spend from your invoice inward, and start measuring live traffic going forward. Both are organised by job and laid out over time — because the point is to watch how the spend changes, not just read a total.

01

Group by job

A workload is the unit every later decision hangs on — a cluster of requests doing the same thing, like one product feature or one agent. When no clean list exists, we rebuild one: the label the app emits, else the caller's identity, else clustering prompts by their fixed structure. Anything that still won't attribute sits in a named bucket, never forced into a bin.

02

Rebuild the spend, month by month

The invoice sets the total finance trusts. Then we split each job into input / output / cached / reasoning — because that mix, not the headline, is what names the fix. Done month by month, it surfaces the early tremors: work climbing to heavier models, ratios drifting, agent loops multiplying, reasoning tokens bloating.

03

Meter live traffic

A gateway and a trace layer — whichever you already run — sit in front of traffic. Every call is stamped with who issued it, where, and the two fields that carry the weight: the task it belongs to (which fuses many calls into one unit of work) and the turn number (which exposes context growth).

Close-out check: the rebuild has to meet the invoice within 5%, or the difference is named — untracked keys, marketplace lines, rounding. We call the read done when tagging covers >90% of traffic, the draft breakdown reconciles within 5%, every job has a named owner, and at least two months of history are in hand to read the trend.

03FOUR LENSES · ELEVEN SIGNALS

What we score, and how.

With the spend finally visible, we check it against four groups of questions — four lenses, eleven signals in all. Each signal gets two readings: a rating on a four-rung ladder (Blind → Visible → Under control → Improving on its own), and a trend showing whether the gap is widening, holding, or closing on its own. The rating isn't the point; the priced gap and its direction are.

Lens A

Can you see it?

Can you see the spend, and see it move?
A1

Traceable spend

Can every dollar on the invoice be traced to a job, a team, and what it bought? Nothing downstream is trustworthy until it can.

How we read itReconcile tagged spend to the invoice; attribute the untagged tail by who owns the keys and by log volume. Many AI-native teams meter nothing below the invoice line, so real coverage is an edge.

0Blindinvoice is the only number
1Visibleprovider / key totals
2Under controlper-job, most traffic
3Improving on its ownper-outcome, self-maintaining

FixesUnlocks every other lens, plus per-customer margin & pricing.

A2

Cost per outcome

Does the setup know the price of one delivered unit — a resolved ticket, a completed task — and is that price tracked and actually used?

How we read itJoin usage to outcome events; where the join is missing, stand up a feature-level proxy and mark it modelled. Set unit cost against the outcome's value to read headroom or alarm.

0Blindnever computed
1Visiblecomputed once, now stale
2Under controlfeature-level proxy, tracked
3Improving on its owntrue unit cost, drives pricing

FixesOutcome-priced billing, invest / retire calls, a margin-aware roadmap.

A3

How fast you catch a spike

How long a cost spike lives before anyone notices — a prompt edit that triples context, a retry storm, an agent that forgets to stop.

How we read itWalk the last quarter for step-changes and, for each, separate when it began from when it was caught. If the record is silent, we stage a controlled spike and see what fires.

0Blindcaught on the invoice
1Visiblecaught in weekly review
2Under controlsame-day automated alerts
3Improving on its ownalerts + cost checks in CI

FixesCost alerts, per-key ceilings, cost-diff checks at deploy.

Lens B

Right-sized models?

Is each job on the cheapest model that still holds up?
B1

Right-sizing

Whether each job runs on the cheapest model that still clears its quality bar — proven by replaying its own traffic, not asserted from a spec sheet. Top-tier to small pricing spans 10–25×; most extraction and classification never needed the top rung.

How we read itTake a representative slice of the job's own prompts and run it down an ordered set of models, cheapest first. Fire each a fixed number of times and grade every result against your own production output — your ground truth, never a benchmark we invented. Stop at the first one that holds. Anything that clears nothing goes to human review with its evidence, not shoved downward.

0Blindone flagship for everything
1Visiblead-hoc tiering, no evidence
2Under controltest-mapped, static
3Improving on its ownre-mapped whenever prices move

FixesStatic remapping — the highest-yield, lowest-effort move — then smaller custom models for narrow high-volume work.

B2

Smart routing

Whether the setup routes by difficulty — cheap model first, escalate only what fails — or sends everything to one tier regardless of how hard the request is.

How we read itFrom the right-sizing tests, isolate jobs where a cheap model passes most of the time but not always — the cascade candidates. Model the blended cost of cheap-first-plus-escalate against the flat-flagship baseline, and price the difference.

0Blindno routing, one tier
1Visiblemanual split by endpoint
2Under controlconfidence cascade in place
3Improving on its ownverifier-gated, self-repricing

FixesConfidence cascades, verifier-gated escalation, provider-fallback routing.

Lens C

What's in each call?

What rides inside each call, and is it bloated?
C1

Context reuse

The shape and reuse of everything sent into the model. A stable prompt prefix can be re-read by the provider at up to a ~90% discount, so what matters is how much of each call is reused rather than freshly billed. Agent workloads run 100:1–300:1 input-to-output because the full context is resent every turn — which makes reuse the lever, not trimming replies.

How we read itRead cached-vs-fresh input straight from provider usage; analyse sampled templates for the reuse ceiling and the cache-breakers — volatile data spliced ahead of stable scaffold, tool lists that reorder, per-session IDs in the system prompt. Price the distance to the ceiling at the cached rate.

0Blindno reuse, profile unknown
1Visibleaccidental hits, known in aggregate
2Under control>60% reuse, structured prefixes
3Improving on its own>85% + a semantic layer where it pays

FixesStatic-prefix restructuring, cache-TTL tuning, an app-level semantic cache.

C2

Output control

Control on the expensive side. Generated tokens price at 5–6× input, and reasoning tokens bill as output — often at a high default nobody set on purpose. The most common quiet overspend we find.

How we read itAudit request settings at every call site in code, joined to trace data. Compare what the model produces to what the downstream code actually consumes — three parsed fields don't need a paragraph wrapped around them. Inspect reasoning settings explicitly; unset or defaulted-high is where it leaks.

0Blindall defaults, prose everywhere
1Visiblecaps on a few endpoints
2Under controlcaps + structure + budgets standard
3Improving on its ownenforced in review & CI

FixesOutput ceilings, structured outputs, per-class reasoning budgets, terser prompts.

C3

Agent-loop cost

Tokens consumed per completed task inside agent loops, where each turn re-sends the whole conversation and tool output so far. It's the fastest-growing line in most setups: a ten-turn task can cost ~50× a single call, and trimming stale context alone recovered ~84% in a published 100-turn evaluation.

How we read itRebuild each loop from grouped traces and measure what fills the window turn by turn — instructions, tool schemas, history, results. Surface the usual offenders: full schemas resent for a two-tool task, multi-KB dumps never read again, no compaction, no stop condition.

0Blindno visibility into loops
1Visiblecall counts known, composition not
2Under controltrimming + compaction in place
3Improving on its owntokens-per-task budgeted & owned

FixesTool-result trimming, scheduled compaction, sub-agent isolation, loop budgets.

Lens D

Discounts & guardrails

Are you paying list price, and are there hard limits?
D1

Discounts claimed

Whether you pay list price when you need not. Discounts already on the table — batch and flex tiers (a slower lane, roughly half price, for work no user is waiting on), caching, volume commitments — frequently go unclaimed. Prices fall ~10×/year at a fixed capability, so a routing choice two quarters old was made against a price sheet that no longer exists.

How we read itSize the eligible discount per class from the inventory, classify each job by how long it can wait, and price async volume stuck on the realtime tier at the flat ~50% recovery. Then find who owns the re-pricing cadence — usually no one.

0Blindlist price on everything
1Visibleone discount class captured
2Under controlmost classes, annual review
3Improving on its ownall classes, quarterly re-pricing

FixesBatch / flex migration, commitment negotiation, a refresh cadence tied to releases.

D2

Wasted spend

Money on calls that returned nothing usable — hard failures, re-parse retries, duplicate submissions, double-completed timeouts.

How we read itRead status codes, retry headers and near-identical prompts in tight windows from gateway logs, then attribute cause — re-parse loops point at missing structured output, 429 bursts at absent backoff, duplicates at missing idempotency. A well-run setup sits in the low single digits.

0Blindnever measured
1Visible>5%, causes unknown
2Under control2–5%, causes known
3Improving on its own<2%, watched continuously

FixesStructured outputs, idempotency keys, jittered backoff, circuit breakers.

D3

Spend limits

Whether spend has hard stops — per-key and per-team ceilings, rate limits, a kill-switch that actually fires — and whether anyone owns the page when one trips.

How we read itAudit gateway and provider config, then prove it — drive a test key into its limit and confirm the block, rather than trusting the setting. Ask who got paged the last time a ceiling tripped, and what happened next.

0Blindno limits anywhere
1Visibleprovider cap only
2Under controlenforced per-key budgets
3Improving on its ownenforced + tested + paged owner

FixesEnforced per-key ceilings, dev / prod key separation, per-job spend targets.

04THE TREND READ

Where your spend is heading.

This is the part a one-time snapshot can't give you, and the reason the whole assessment runs over time. Your spend is never still — model line-ups change, agent features spread, prompts grow, providers cut prices. A number that looked fine two quarters ago can be a live leak today with nothing obviously broken. We track each of these movements and project it forward.

Climb to heavier models

Spend drifting up the model stack. Usually one team adopts a new reasoning model and it seeps into jobs that never needed it.

Agent share

The slice of the bill sitting inside agent loops. A flat total can hide an agent share quietly doubling under it.

Context growth

The per-turn growth curve over time. Steepening month over month means history-replay cost is compounding.

Reasoning share

Thinking tokens as a rising fraction of output — the quiet drift as more work lands on reasoning-capable models.

Reply-length creep

Typical reply length per job over time, held against what the consuming code actually reads.

Waste trend

Waste as a moving ratio, not a point — retry storms and duplicates tend to grow with traffic until someone trips over them.

Unit-cost curve

Cost per outcome against volume growth. Rising unit cost on rising volume is the alarm finance actually watches for.

Price-sheet risk

Provider repricing and new releases you haven't reacted to — the outside axis moving whether or not anyone looks.

The do-nothing curve: from the measured motion we extend the bill two-to-four quarters out on the assumption that nobody intervenes — always modelled with a range, never a single line. What we find again and again: the total can sit flat while the mix underneath churns — and "just move everything to the newest model" often lifts the bill rather than lowering it. The value was never in the total; it's in the per-job distribution.

05PRICING THE FIXES

Putting a dollar figure on each fix.

A finding nobody has priced is just an opinion, so every gap and trend signal is turned into money the same way, each line carrying its evidence tag. We price each fix twice — against what you spend today, and against the do-nothing forecast you're already heading toward — because that second, growing number is what shows the real cost of waiting.

Do-nothing

Where it lands untouched

Where the trend read puts the bill if nobody acts. Every saving below is measured against this line, not just against today.

modelled · range
Right model

Test-proven

Every job on the cheapest model it actually cleared in the tests. This curve is measured, not estimated.

measured
Model + payload + discounts

Everything the profile unlocks

Fixes applied in order — route, then cache, then batch, then trim — because they stack, they don't add.

stacked

Rule: fixes stack, they don't add. Route a job to a model 10× cheaper and you've shrunk the base that caching then discounts — so they're applied in sequence, never summed, or the total lies. Every recommendation also carries a build estimate; anything paying back beyond two quarters is roadmap, not a quick win, and it's labelled as such.

06SHIP & VERIFY

What we ship during the run.

Most reviews end with a list of recommendations. Where we find easy, low-risk wins during the run, we take them — implement them and verify the saving on your own traffic — so the readout can open with results already proven, not just projected. If nothing is safe to ship yet, we say so rather than force it.

If a rebuild is worth it — the order it goes in
1Evals first — nothing routes safely without them.
2Routing & cascades, built on top of the evals.
3Agent context engineering and semantic caching.
4Smaller custom models last, for stable high-volume tasks.
What you keep running it
ProofBefore-and-after on the same traffic, with a holdout where volume allows. A saving is only claimed once it shows up in the gateway data — nothing booked on a promise.
The targetA cost-per-outcome model in your hands: today, after the quick wins, after the rebuild — plotted against volume growth and against the do-nothing curve.
Operating modelDashboards, a budget policy with named owners, a monthly review. Two standing rules: every provider price change triggers a re-routing pass, and the trend axes are watched so slippage is caught in-month, not on the invoice.
07THE READOUT

The final readout.

It all lands in a single ninety-minute session, with engineering and finance in the same room — because the real value is a conversation those two sides can finally have over one set of numbers. The running order:

0–15
The readWhat you spend, where it goes, and how the shape has moved across the window.
15–35
The scorecardWhere you stand across the four lenses — every rating and trend backed by the query behind it.
35–50
The trend & do-nothing forecastWhere the spend is heading, and what the bill becomes if nobody acts.
50–65
The priced fixesWhat each gap is worth — evidence-tagged, three curves, a floor and a likely number.
65–75
What already shippedThe wins landed during the run, proven in your own gateway data.
75–85
Rebuild & targetThe sequenced plan and the curve you could be on.
85–90
Phase two, if there is oneOutcome-linked commercials, anchored on the floor.

And if the read comes back lean and the curve comes back flat, this meeting is ten minutes and an all-clear. We'd rather hand back a clean bill of health than talk you into a rebuild you don't need.