Three to four weeks on one named agent workflow — Diagnose inventories every system it reads structure from and the connector each needs, and scores all seven stages with evidence; Blueprint sequences the Core Model we are willing to put in writing.
Published · Revised
$32,5003 to 4 weeks — Diagnose $15,000 plus Blueprint $17,500
Briefing on the seven-stage Context Supply Chain and maturity model, applied to the named workflow
Connector assessment: inventory of every schema-bearing system the workflow touches (data and content platforms, warehouse, API, ontology), the format each speaks, the connector or standard each needs, and the direction of flow
L0–L5 maturity scorecard with evidence notes and gap analysis
The Core Model, sequenced: which schemas import first, how they harmonize across systems, the conformance contract the boundaries will enforce, and the agents that read from it
Prioritized remediation roadmap with the first implementation quarter scoped
Detail
The engagement in full
What it is
The Core Model Blueprint is where directional judgment becomes measurement, on one named agent workflow. Where the Diagnostic Briefing places you on the maturity scale from a half day of structured discussion, the Blueprint earns the placement: content is sampled, retrieval is tested against a ground-truth question set, semantic drift is checked across the systems the workflow reads, and every score on the L0–L5 card carries an evidence note saying what was examined and what it showed.
Diagnose is, concretely, a connector assessment. Every system the workflow reads structure or meaning from — data and content platforms, the warehouse, APIs, ontologies, the wiki that has quietly become a source of truth — is inventoried with the format it speaks and the connector or public standard it would need to feed one governed model. Blueprint then sequences the Core Model itself: what imports first, how the schemas harmonize, what conformance contract the boundaries enforce, and which agents read from it through the MCP endpoint.
It splits into two separately priced phases, and the split is real, not decorative. Diagnose ($15,000, roughly two weeks) establishes what is true. Blueprint ($17,500, roughly two weeks) decides what to do about it. You can commission the first and decline the second — the diagnosis is a complete deliverable on its own, and some clients take it to their own engineering organization and stop there. That is a supported outcome, not a failed sale.
Capacity is genuinely limited to three engagements per quarter. This is a delivery constraint, not a scarcity device: the assessment is run by the practice lead, not staffed out, because a scorecard is only as good as the judgment behind the evidence notes.
Scope is one named agent workflow, chosen with you at kickoff — usually the one the Audit or the interview map pointed at. That is a tightening, on purpose: seven stages measured for one workflow yields a Core Model your engineers can build next quarter; seven stages surveyed across an estate yields a report.
Who it is for
Enterprises with multiple agent pilots on multiple retrieval stacks and no shared answer for why quality varies between them — and one workflow they would fix first
Regulated organizations that must show what an agent read, where it came from, and what checked it — and currently cannot
Teams that funded diagnosis after a briefing, and want the implementation decision made on measurement rather than momentum
Engineering leaders who intend to build the fixes themselves and want the diagnosis and sequencing done by someone who has seen the failure patterns before
What happens: Diagnose (weeks 1–2)
Each of the seven stages gets its own examination for the named workflow, with method matched to stage:
Sources — inventory of what content exists, who owns it, and freshness sampling: how much of what agents can reach has been superseded, and would anyone know
Structure — a content sample audited for machine-legible shape: what has schema, what has convention, what has neither
Semantics — drift check across systems: does “customer” (or “policy”, or “asset”) mean one thing everywhere it appears, and where the meanings fork, does anything record the fork
Validation — inventory of what checks content before an agent reads it; usually the shortest chapter, which is itself the finding
Retrieval — fitness testing against a ground-truth question set built with your domain owners: what should have been returned versus what was
Agent Operations — observability review: when an agent misbehaves in production, what signal exists, who sees it, and how fast
Learning Loop — the uncomfortable question, asked with evidence: when did a failure last change the system, and can you point at the change
Output: the L0–L5 maturity scorecard, one score per stage, each with evidence notes and a gap analysis. Not an average — a chain is as strong as its weakest defended stage, and the card says which stage that is.
What happens: Blueprint (weeks 3–4)
The blueprint answers three questions the diagnosis raises.
What level does each stage actually need?
Not L5 everywhere — recommending maximum maturity across the board would be malpractice by spreadsheet. A customer-facing regulated workflow and an internal drafting assistant justify different targets, and the blueprint sets targets per stage per use case, with the reasoning shown.
In what order?
Sequencing by dependency and risk. Retrieval built on unstructured sources rebuilds itself in a year; the blueprint sequences so no quarter's work is torn up by the next quarter's.
At what cost?
The first implementation quarter scoped against the fixed package on the next rung, and any expansion quarters priced by scope — so what you carry into a funding conversation is a costed plan, not an aspiration.
And one more, made explicitly: a go or no-go recommendation, including “do not implement yet”. If the honest read is that your sources need an ownership model before any build is worth funding, the blueprint says so, in writing, at the cost of the larger engagement it forgoes.
What you walk away with
The seven-stage current-state map of the workflow’s context supply chain
The L0–L5 scorecard with per-stage evidence notes — procurement-ready, and reusable in vendor evaluations far beyond this engagement
A prioritized remediation roadmap, sequenced by dependency and risk, with the first implementation quarter costed
A written go or no-go recommendation
Everything owned by you outright — executable by your engineers, by another partner, or by us; the ladder is priced so that this is a real choice
The situations this exists for
Three pilots, three stacks, one question
A large organization ran agent pilots in three business units; one performs well, two do not, and each team has a theory involving another team's systems. The Blueprint replaces three theories with one measured answer on the workflow they pick first — typically the difference lives in the stages, not the models, and the scorecard shows exactly where.
The provenance requirement
A compliance function asks the AI team to show what an agent read before it answered, and the honest current answer is a retrieval log nobody can interpret. The validation and agent-operations chapters of the diagnosis become the gap register the compliance program is built against.
The build-it-ourselves plan
An engineering organization fully intends to do the implementation internally — it wants the diagnosis, targets, and sequencing from someone who has watched these chains fail before. Diagnose-only, blueprint declined or taken and self-executed, is a designed-for path, and the deliverable ownership terms exist precisely for it.
What it is not
No code is shipped and no system is changed. Access is read-only: telemetry, a content sample, and time with the people who own the sources.
It is not an audit of your models. When the model genuinely is the limit, the diagnosis says so — but in most estates we examine, model limits are the smaller share of the failure surface, and the evidence notes let you check that claim rather than take it.
The blueprint is not a proposal in costume. It is priced, sequenced engineering scope that any competent team can execute. If it only made sense when we execute it, it would be a sales document, and it would be worth less to you.
Pricing
This engagement on the ladder
Rung 3
Core Model Blueprint
$32,5003 to 4 weeks — Diagnose $15,000 plus Blueprint $17,500
Enterprises ready to diagnose one workflow's connectors and blueprint its Core Model.
One named agent workflow, diagnosed across seven stages
Diagnose: the connector assessment
Blueprint: the Core Model, sequenced
Conformance and efficiency protocols review
A go or no-go recommendation — including 'do not implement yet'
The interview, then a written fixed-scope fixed-fee proposal. We recommend starting at the Agent Context Cost & Failure Audit — it is approvable without a steering committee. Nothing requires you to enter at the bottom of the ladder — a Diagnostic Briefing within ten days of the interview is common when the funding decision is contested.
What system access do you need?
Assessment work is read-only: usage telemetry, a content sample, and time with the people who own the sources. Implementation access is scoped explicitly at the start of each quarter and is limited to the systems in that quarter's scope.
Who owns the deliverables?
You do. Reports, blueprints, runbooks and code produced in an engagement are yours outright. A blueprint you commission from us can be executed by your own engineers or by another partner — that is a deliberate property of how the ladder is priced, not a concession.
Can we do Diagnose only?
Yes, and clients do. The diagnosis is complete on its own — scorecard, evidence notes, gap analysis. The blueprint adds targets, sequencing, and costed scope; if your own organization does that work well, take the diagnosis and run.
How is the ground-truth question set built?
With your domain owners, in week one — questions where the correct source material is known and agreed before retrieval is tested against it. It stays with you afterward, and it is the seed of the evaluation harness the Implementation rung builds out.
When two systems disagree on the same metric, how is it resolved and can I reproduce yesterday's answer?
The disagreement is resolved as a mapping in the Core Model, approved by the owner of each schema, and recorded in the change log with who changed what and when. Agents read the model through the MCP endpoint, so an answer traces to the model version it was read from. Version history is a Team and Enterprise feature on coremodels.io; with it in place, yesterday's version is read back, and the answer with it.
What evidence do you hand my auditors when an agent answer becomes a finding?
The change log, showing who changed what and when. The exported conformance contract in force at the time. The mapping the answer depended on, and who approved it. And, on Managed Operations' Ops + Eval, the incident attribution: whether the failure was in the context, the model or the application.
Every refuse or warn decision is a record with named fields: who approved the mapping and when, the timestamped freeze record tied to the contested mapping id, who ran the check, the agent or workflow id that asked and a hash of the input it sent, the UTC timestamp, the Core Model project with its schema version and the element version in force, the ShEx shape and the mapping it failed, the check that fired and its reason code, and the answer that was refused or warned. Records are retained for the period set in your engagement letter, alongside your other audit artefacts. Your named control owner, and anyone they authorize, can export the pack as JSON at any time.
What is deterministic is the decision, not the prose: for a given contract version and mapping, the same input produces the same refuse or the same warn every time. Your evals assert that decision, not the prose. A generated answer is not reproducible word for word, which is why the boundary, not the model, is what you audit. The engagement maps each boundary check to a control in the framework you already report against — BCBS 239, Solvency II, IFRS 17 and 21 CFR Part 11 are the shapes it takes most often.
When an agent answer becomes a finding, that pack — the refuse or warn record, the approved mapping and the contract version in force — is what your auditors receive, and because the decision is deterministic they can re-run it without us in the room.