Three to four weeks of measurement across all seven stages of your context supply chain — ending in an evidenced L0–L5 scorecard, a sequenced blueprint, and a go or no-go recommendation we are willing to make in writing.
Published · Revised · ARIES-PROD-003
$32,5003 to 4 weeks — Diagnose $15,000 plus Blueprint $17,500
Duration
3 to 4 weeks
Total fee
$32,500
Diagnose phase
$15,000
Blueprint phase
$17,500
Capacity
3 engagements per quarter
Deliverables
Executive briefing on the seven-stage Context Supply Chain and maturity model
Current-state map of sources, structure, semantics, validation, retrieval, agent ops, and learning
L0–L5 maturity scorecard with evidence notes and gap analysis
Prioritized remediation roadmap with recommended engagement scope and sequence
Detail
The engagement in full
What it is
The Assessment is where directional judgment becomes measurement. Where the Executive Briefing places you on the maturity scale from a half day of structured discussion, the Assessment earns the placement: content is sampled, retrieval is tested against a ground-truth question set, semantic drift is checked across systems, and every score on the L0–L5 card carries an evidence note saying what was examined and what it showed.
It splits into two separately priced phases, and the split is real, not decorative. Diagnose ($15,000, roughly two weeks) establishes what is true. Blueprint ($17,500, roughly two weeks) decides what to do about it. You can commission the first and decline the second — the diagnosis is a complete deliverable on its own, and some clients take it to their own engineering organization and stop there. That is a supported outcome, not a failed sale.
Capacity is genuinely limited to three engagements per quarter. This is a delivery constraint, not a scarcity device: the assessment is run by the practice lead, not staffed out, because a scorecard is only as good as the judgment behind the evidence notes.
Who it is for
Enterprises with multiple agent pilots on multiple retrieval stacks and no shared answer for why quality varies between them
Regulated organizations that must show what an agent read, where it came from, and what checked it — and currently cannot
Teams that funded diagnosis after a briefing, and want the implementation decision made on measurement rather than momentum
Engineering leaders who intend to build the fixes themselves and want the diagnosis and sequencing done by someone who has seen the failure patterns before
What happens: Diagnose (weeks 1–2)
Each of the seven stages gets its own examination, with method matched to stage:
Sources — inventory of what content exists, who owns it, and freshness sampling: how much of what agents can reach has been superseded, and would anyone know
Structure — a content sample audited for machine-legible shape: what has schema, what has convention, what has neither
Semantics — drift check across systems: does “customer” (or “policy”, or “asset”) mean one thing everywhere it appears, and where the meanings fork, does anything record the fork
Validation — inventory of what checks content before an agent reads it; usually the shortest chapter, which is itself the finding
Retrieval — fitness testing against a ground-truth question set built with your domain owners: what should have been returned versus what was
Agent Operations — observability review: when an agent misbehaves in production, what signal exists, who sees it, and how fast
Learning Loop — the uncomfortable question, asked with evidence: when did a failure last change the system, and can you point at the change
Output: the L0–L5 maturity scorecard, one score per stage, each with evidence notes and a gap analysis. Not an average — a chain is as strong as its weakest defended stage, and the card says which stage that is.
What happens: Blueprint (weeks 3–4)
The blueprint answers three questions the diagnosis raises.
What level does each stage actually need?
Not L5 everywhere — recommending maximum maturity across the board would be malpractice by spreadsheet. A customer-facing regulated workflow and an internal drafting assistant justify different targets, and the blueprint sets targets per stage per use case, with the reasoning shown.
In what order?
Sequencing by dependency and risk. Retrieval built on unstructured sources rebuilds itself in a year; the blueprint sequences so no quarter's work is torn up by the next quarter's.
At what cost?
Implementation quarters scoped and priced against the fee range on the next rung — so what you carry into a funding conversation is a costed plan, not an aspiration.
And one more, made explicitly: a go or no-go recommendation, including “do not implement yet”. If the honest read is that your sources need an ownership model before any build is worth funding, the blueprint says so, in writing, at the cost of the larger engagement it forgoes.
What you walk away with
The seven-stage current-state map of your estate
The L0–L5 scorecard with per-stage evidence notes — procurement-ready, and reusable in vendor evaluations far beyond this engagement
A prioritized remediation roadmap, sequenced by dependency and risk, with implementation quarters costed
A written go or no-go recommendation
Everything owned by you outright — executable by your engineers, by another partner, or by us; the ladder is priced so that this is a real choice
The situations this exists for
Three pilots, three stacks, one question
A large organization ran agent pilots in three business units; one performs well, two do not, and each team has a theory involving another team's systems. The assessment replaces three theories with one measured answer — typically the difference lives in the stages, not the models, and the scorecard shows exactly where.
The provenance requirement
A compliance function asks the AI team to show what an agent read before it answered, and the honest current answer is a retrieval log nobody can interpret. The validation and agent-operations chapters of the diagnosis become the gap register the compliance program is built against.
The build-it-ourselves plan
An engineering organization fully intends to do the implementation internally — it wants the diagnosis, targets, and sequencing from someone who has watched these chains fail before. Diagnose-only, blueprint declined or taken and self-executed, is a designed-for path, and the deliverable ownership terms exist precisely for it.
What it is not
No code is shipped and no system is changed. Assessment access is read-only: telemetry, a content sample, and time with the people who own the sources.
It is not an audit of your models. When the model genuinely is the limit, the diagnosis says so — but in most estates we examine, model limits are the smaller share of the failure surface, and the evidence notes let you check that claim rather than take it.
The blueprint is not a proposal in costume. It is priced, sequenced engineering scope that any competent team can execute. If it only made sense when we execute it, it would be a sales document, and it would be worth less to you.
Pricing
This engagement on the ladder
Rung 3
Context Supply Chain Assessment
$32,5003 to 4 weeks — Diagnose $15,000 plus Blueprint $17,500
Enterprises ready to diagnose and blueprint the full chain.
Capacity-limited to 3 engagements per quarter
Diagnose across seven stages
Blueprint for target maturity
Conformance and efficiency protocols review
Phased implementation options
Limited to three per quarter
A go or no-go recommendation — including 'do not implement yet'
A scoping call, then a written fixed-scope fixed-fee proposal. Most buyers start at the Token Economics Audit because it is approvable without a steering committee. Nothing requires you to enter at the bottom of the ladder — an executive briefing first is common when the funding decision is contested.
What system access do you need?
Assessment work is read-only: usage telemetry, a content sample, and time with the people who own the sources. Implementation access is scoped explicitly at the start of each quarter and is limited to the systems in that quarter's scope.
Who owns the deliverables?
You do. Reports, blueprints, runbooks and code produced in an engagement are yours outright. A blueprint you commission from us can be executed by your own engineers or by another partner — that is a deliberate property of how the ladder is priced, not a concession.
Can we do Diagnose only?
Yes, and clients do. The diagnosis is complete on its own — scorecard, evidence notes, gap analysis. The blueprint adds targets, sequencing, and costed scope; if your own organization does that work well, take the diagnosis and run.
How is the ground-truth question set built?
With your domain owners, in week one — questions where the correct source material is known and agreed before retrieval is tested against it. It stays with you afterward, and it is the seed of the evaluation harness the Implementation rung builds out.