Two weeks, one named agent workflow, two questions: what is each unit of its work actually costing you, and where does it fail because the context it was handed was wrong — the wrong tool, the wrong field, a definition that went stale?
Published · Revised
$12,0002-week fixed engagement — one named agent workflow
Duration
2 weeks
Scope
One named agent workflow
Fee
$12,000
Output
CFO-ready written report
Access needed
Usage telemetry and agent traces, read only
Deliverables
Workload inventory for the named agent workflow with measured cost attribution
Wrong-context failure register: wrong tool, wrong field, stale definition — each failure traced to the context the agent was handed
The share of spend attributable to schemas that do not agree — what a Core Model would remove
CFO-ready report with methodology stated in full
Prioritized remediation list with estimated recovery
Detail
The engagement in full
What it is
The Agent Context Cost & Failure Audit takes one named agent workflow — resolving a ticket, drafting a document, completing a workflow step — and measures two things about it: what each unit of its work costs, and where it fails because the context it was handed was wrong.
One line of the report is new since 2026-09: the share of spend that pays for context your systems could have agreed on once — structure re-derived on every call, mappings guessed by the model, the same definition fetched from three places. That line is the business case for a governed Core Model stated in your own numbers, and the Core Model Blueprint is where it becomes a plan.
The cost side is token economics. Cost per million tokens is the number on the invoice. It is also the wrong number. Token prices have fallen for two years while agent bills have risen, because the waste lives in the ratio: how many tokens a unit of work consumes, and how many of those tokens the model never needed. That ratio is invisible on the invoice and measurable in your telemetry, and the audit exists to measure it.
The failure side is wrong context. An agent that calls the wrong tool, reads the wrong field, or trusts a definition that went stale three releases ago does not look like a context failure in the logs — it looks like a model being unreliable. We trace each failure in the workflow back to the context the agent was handed and count the failures by mode: wrong tool, wrong field, stale definition.
2 weeks
From kickoff to report
Scope is fixed at kickoff: one named agent workflow. The fee does not move.
This is the entry diagnostic — deliberately small enough to approve without a steering committee, and concrete enough that the finding stands on its own whether or not you do anything further with us. The report states its methodology in full, including what it could not measure. A finance reviewer should be able to reproduce the arithmetic without calling us.
Who it is for
The finance leader whose model spend doubled while measured output did not
The AI or platform lead whose agent is unreliable in production and cannot yet say whether the model or the context is at fault — and who needs a defensible baseline before committing to a larger program
Procurement teams sizing an enterprise commitment to a model provider and unwilling to anchor it on last quarter's unexamined usage
Anyone who has been asked what one agent transaction costs us, in a meeting, and had to answer with the invoice total divided by a guess
What happens, week by week
Week 1 — Inventory and definition
Read-only ingestion of your usage telemetry and the workflow’s agent traces. We build the inventory for the named workflow — every step that spends tokens, and every tool, field and definition it reaches for — and agree the unit-of-work definitions with your team, because a unit chosen unilaterally by an auditor is a unit finance will reject.
Week 2 — Attribution and analysis
Cost attributed per step and per unit. Waste patterns identified and quantified. Failures traced back to the context the agent was handed and counted by mode. The report is drafted, reviewed with your team for factual accuracy — not for conclusions, those are ours — and delivered with a readout session.
The waste patterns we find, and keep finding
Retrieval overfetch — top-k padding pushed into context just in case, paid for on every call, read by the model almost never
Retry loops — failed calls re-sent with full context attached, so one failure costs three successes
Redundant fetching — step four of a pipeline retrieving what step two already retrieved, because no stage knows what another stage holds
Unbounded history — conversation context growing linearly with session length, so the hundredth turn costs an order of magnitude more than the first
Boilerplate ballast — instruction blocks and few-shot examples carried on every call, long after the workflow stopped needing them
None of these show on an invoice. All of them show in telemetry.
The wrong-context failures we find
Wrong tool — the agent has a dozen tools and picks by description; two descriptions overlap, and the wrong one wins often enough to show up in the numbers
Wrong field — two systems each have a field called status and they mean different things; the agent reads the one it was shown first, and the answer is confidently wrong
Stale definition — a schema changed, the prompt or the retrieval index did not, and the agent is acting on a definition nobody has used since spring
None of these show in a model benchmark either. All of them show in the traces once the failure is traced to the context rather than blamed on the model — and each is fixable in the schema once instead of in the prompt forever.
What you walk away with
A workload inventory with measured cost attribution — what each agent workflow costs per unit of work
A waste analysis quantifying each pattern found in your estate
A wrong-context failure register for the workflow — each failure traced to its mode, with the context fix that would have prevented it
A prioritized remediation list with modeled recovery, presented as ranges with assumptions stated — modeled from your own telemetry, and not guarantees
A CFO-ready report with methodology in full, built to survive a finance review
A clear go or no-go for the Core Model Blueprint — including “your spend is fine; your problem is elsewhere”
The situations this exists for
The bill that doubled
Usage grew modestly; spend more than doubled. Nobody changed anything on purpose. Two weeks of telemetry work typically finds the growth living in two or three of the waste patterns above — and finds which workflow they live in, which is what makes remediation a task instead of a debate.
The commitment decision
An enterprise agreement with a model provider is on the table, sized from current consumption. If a third of current consumption is buying nothing, the commitment is mis-sized by a third. The audit produces the baseline before the signature, not after.
The efficiency claim
An internal team reports a large cost reduction from a caching change and wants to scale it estate-wide. The audit verifies the claim against measured units of work before the pattern is replicated — because scaling a measurement error scales the error.
The unreliable agent
An agent that scores well in evaluation misbehaves in production on one workflow, and the team has spent a month tuning prompts. Tracing the failures back to context usually finds a handful of wrong-field and stale-definition modes — fixable in the schema once, rather than in the prompt forever.
What it is not
It is not an assessment of the whole estate. The audit covers one named workflow: it tells you what spend buys nothing and where the context was wrong; it does not score all seven stages or sequence a Core Model. That is the Blueprint’s territory.
It is not a procurement negotiation service. The report will strengthen your negotiating position; we do not sit at that table.
It does not promise savings. When the honest finding is that your estate is efficient, that is the finding you get, and it is worth exactly as much.
Pricing
This engagement on the ladder
Rung 2
Agent Context Cost & Failure Audit
$12,0002-week fixed engagement — one named agent workflow
Teams needing a fast, procurement-ready commercial entry on one agent workflow.
Two weeks on one named agent workflow
Cost per unit of work from your telemetry
Wrong-context failures: wrong tool, wrong field, stale definition
The interview, then a written fixed-scope fixed-fee proposal. We recommend starting at the Agent Context Cost & Failure Audit — it is approvable without a steering committee. Nothing requires you to enter at the bottom of the ladder — a Diagnostic Briefing within ten days of the interview is common when the funding decision is contested.
What system access do you need?
Assessment work is read-only: usage telemetry, a content sample, and time with the people who own the sources. Implementation access is scoped explicitly at the start of each quarter and is limited to the systems in that quarter's scope.
Who owns the deliverables?
You do. Reports, blueprints, runbooks and code produced in an engagement are yours outright. A blueprint you commission from us can be executed by your own engineers or by another partner — that is a deliberate property of how the ladder is priced, not a concession.
What telemetry do you need, exactly?
API usage logs with per-call token counts, workflow or session identifiers, the named workflow’s agent traces (tool calls and the fields they read), and whatever attribution you already have (team, application, environment). Read-only, and a gap in your telemetry becomes a stated limitation in the report rather than a silent assumption.
What if we do not have per-workflow attribution?
Then building enough of it to attribute cost is week one's work, and the report says which numbers are measured and which are apportioned. The absence of attribution is itself a finding — it usually means nobody can currently answer the cost question at all.