Monthly ownership of your Core Model and the supply chain around it, in two SKUs. Care keeps the model true — new connectors as systems change, drift caught by the change log. Ops + Eval adds the evaluation harness run on schedule, incident attribution, and on-call for the named agent workflows. Both report quarterly in ROKA terms: what changed, what it cost, what it returned.
Published · Revised
$9,000 – $18,000Per month — Care $9,000 · Ops + Eval $18,000; quarterly rolling; CoreModels licence included
Cadence
Monthly
Fee
Care $9,000 · Ops + Eval $18,000 per month
Reporting
Quarterly ROKA report
Commitment
Quarterly, rolling
Licence
CoreModels licence included for the engagement term
The CoreModels licence for the engagement term — licensed to your team and readable by your MCP tools on your own agent API key; Managed Operations keeps it current
Care: managed operations coverage for agreed stages of the Context Supply Chain
Care: Core Model stewardship — connectors added as systems change, mappings and terminology kept in step, drift caught in the change log
Care: ongoing monitoring of context health and source freshness
Ops + Eval: the evaluation harness maintained and run on schedule, retrieval fitness re-tested against ground truth
Ops + Eval: incident attribution by layer and on-call for the named agent workflows
Ops + Eval: learning-loop updates that feed validated operational insight back into the model, its sources and its structure
Quarterly ROKA report and scheduled reviews of maturity posture with change recommendations
Detail
The engagement in full
What it is
Stage seven, the learning loop, is the stage that quietly stops running. Evaluation suites go stale. Sources drift. A policy library gets superseded and the retrieval index does not hear about it. Nobody notices until an agent answers confidently from material that was replaced a year ago — and by then the question of who owns context health has the worst possible answer, which is everyone, meaning no one.
Operationally, most of the month is spent keeping the Core Model true: a system gets added and needs a connector, a team changes a schema and the change log shows the drift before an agent reads it, an incident traces back to a mapping and the mapping is fixed in the model rather than in a prompt. The learning loop is not a report; it is the list of changes made to the model, with reasons.
Managed Operations exists because that stage needs a named owner. It is a monthly engagement, on a quarterly rolling commitment, in two SKUs. Care, at $9,000 a month, keeps the Core Model true: source freshness monitored, new connectors added as systems change, mappings and terminology kept in step, drift caught in the change log before an agent reads it. Ops + Eval, at $18,000 a month, adds what Care does not: the evaluation harness run on schedule against ground truth, every production incident attributed to its layer — context, model, or application — and on-call for the named agent workflows.
It is deliberately the smallest rung by monthly cost and the longest by relationship, because context health is not a project. It is operated infrastructure, and the alternative to operating it is watching a well-built chain drift back to ad hoc over four quarters — a decay we have watched often enough to price the prevention.
Care, and what Ops + Eval adds
What each SKU includes
Care
Ops + Eval
Monthly fee
$9,000
$18,000
CoreModels licence for the engagement term
Included
Included
Core Model stewardship — connectors, mappings, drift in the change log
Included
Included
Source freshness monitoring
Included
Included
Quarterly ROKA report
Included
Included
Evaluation harness run on schedule against ground truth
—
Included
Incident attribution by layer: context, model, application
—
Included
On-call for the named agent workflows
—
Included
Learning-loop execution: failures fed back into the model
—
Included
Who it is for
Teams whose go-live went well and who intend for that sentence to still be true in eighteen months
Organizations whose content estate changes weekly — product catalogs, policy libraries, clinical documentation — where the chain must keep pace with the sources or become a liability with a search box
Internal platform teams that own the agent runtime but want the learning loop externally owned and externally accountable
Anyone who has discovered that keeping the eval suite updated ourselves is the most confidently made and least kept promise in production AI
What happens, monthly
Care — Core Model stewardship: a system gets added and needs a connector; a team changes a schema and the change log shows the drift before an agent reads it; mappings and terminology are kept in step
Care — source freshness monitoring: sampling the content estate for supersession and drift, so stale material is found by us before it is found by an agent
Care — backlog and cadence: a maintained continuous-improvement backlog, groomed monthly, drawn on quarterly
Ops + Eval — retrieval fitness re-testing: the ground-truth question set run against production retrieval on a schedule, so quality regression is a detected event with a date, not a slow suspicion
Ops + Eval — evaluation suite maintenance: new failure modes become new test cases; the suite grows with the estate instead of fossilizing at go-live
Ops + Eval — incident attribution and on-call: production behavior signals triaged, each incident attributed to its actual layer — context, model, or application — with an on-call owner for the named agent workflows. The attribution matters, because the three have different owners and different fixes
Ops + Eval — learning-loop execution: validated operational findings routed back into sources and structure, with the change recorded. The loop’s output is changes made, not observations filed
The quarterly ROKA report
ROKA — Return on Knowledge Assets — is the quarterly statement of account: what changed in the chain, what it cost, and what it returned, in terms a business reader can check. Freshness posture by source. Connectors added and drift caught in the change log. On Ops + Eval, the retrieval fitness trend against the ground-truth set and incidents by layer with time-to-attribution. Changes shipped through the learning loop, each traceable from the failure that prompted it to the change that resolved it.
It is written so the executive who funded the chain can see whether it is still earning its keep — and so that if the honest trend line is flat, the report says flat, and the review meeting is about why.
What you walk away with, quarter after quarter
A context supply chain whose health is measured on a cadence instead of assumed
An evaluation harness that is alive — growing with the estate, run on schedule, trusted enough to gate changes (Ops + Eval)
Incident attribution that ends the is-it-the-model-or-is-it-us argument with evidence, each time — and an on-call owner for the named agent workflows (Ops + Eval)
A ROKA report your finance and leadership can read without translation
One accountable partner for long-running context work — a named owner for the stage that otherwise has none
The situations this exists for
The eighteen-month drift
A chain built well at go-live degrades on no particular day. Sources turn over, the eval suite stops being run, and two years in, answer quality is visibly worse with no visible cause. Managed operations is priced against exactly this decay curve — the monitoring exists so the drift is caught in month two, not month eighteen.
The estate that will not sit still
A commerce platform's catalog changes daily; an insurer's policy library changes with every regulatory cycle. For these estates the chain is never done, and the question is only whether keeping it current is someone's job or everyone's assumption. This rung makes it a job.
The runtime-here, loop-there split
A strong internal team runs the agents and wants to keep running them — but wants the learning loop owned by a party whose quarterly report will say flat when the trend is flat. External accountability for stage seven, internal ownership of everything else, is a clean and common split, and it is usually Ops + Eval: the harness and the incident attribution are exactly what the internal team wants externally owned.
What it is not
It is not a helpdesk or a NOC for your AI stack. Coverage is the agreed stages of the context supply chain, and on-call on Ops + Eval is for the named agent workflows — model provider incidents, application bugs, and infrastructure outages have their own owners, and the attribution work exists partly to route things to them quickly.
It is not a substitute for source ownership. Your people still own the content and make the editorial calls; we detect, attribute, route, and verify. A managed loop with no engaged source owners degrades into a very well-documented backlog, and we will say so if we see it happening.
It is not open-ended. Quarterly rolling means each quarter is a renewal decision made against a ROKA report — which is exactly the discipline we would want if we were the buyer.
Pricing
This engagement on the ladder
Rung 5
Managed Operations
$9,000 – $18,000Per month — Care $9,000 · Ops + Eval $18,000; quarterly rolling; CoreModels licence included
Teams keeping the Core Model and the chain around it true after go-live.
Care ($9,000): source freshness and Core Model drift monitored monthly
Care: new connectors as systems change, mappings kept in step
Care: quarterly ROKA report
Ops + Eval ($18,000) adds: the evaluation harness run on schedule against ground truth
The interview, then a written fixed-scope fixed-fee proposal. We recommend starting at the Agent Context Cost & Failure Audit — it is approvable without a steering committee. Nothing requires you to enter at the bottom of the ladder — a Diagnostic Briefing within ten days of the interview is common when the funding decision is contested.
What system access do you need?
Assessment work is read-only: usage telemetry, a content sample, and time with the people who own the sources. Implementation access is scoped explicitly at the start of each quarter and is limited to the systems in that quarter's scope.
Who owns the deliverables?
You do. Reports, blueprints, runbooks and code produced in an engagement are yours outright. A blueprint you commission from us can be executed by your own engineers or by another partner — that is a deliberate property of how the ladder is priced, not a concession.
Care or Ops + Eval: what decides which we need?
Whether the evaluation harness, incident attribution and on-call have another owner. Care, at $9,000 a month, keeps the Core Model true — connectors, mappings, drift caught in the change log. Ops + Eval, at $18,000, adds the harness run on schedule against ground truth, every incident attributed to its layer, and on-call for the named agent workflows. It is set at each quarterly renewal against the actual estate, not projected once and forgotten.
Do you need production write access?
Minimal and scoped. Monitoring and testing are largely read-only; learning-loop changes to sources and structure go through your own change process with your owners approving. We would rather the loop be slightly slower than be a party with unaudited write access to your content estate.
When two systems disagree on the same metric, how is it resolved and can I reproduce yesterday's answer?
The disagreement is resolved as a mapping in the Core Model, approved by the owner of each schema, and recorded in the change log with who changed what and when. Agents read the model through the MCP endpoint, so an answer traces to the model version it was read from. Version history is a Team and Enterprise feature on coremodels.io; with it in place, yesterday's version is read back, and the answer with it.
What evidence do you hand my auditors when an agent answer becomes a finding?
The change log, showing who changed what and when. The exported conformance contract in force at the time. The mapping the answer depended on, and who approved it. And, on Managed Operations' Ops + Eval, the incident attribution: whether the failure was in the context, the model or the application.
Every refuse or warn decision is a record with named fields: who approved the mapping and when, the timestamped freeze record tied to the contested mapping id, who ran the check, the agent or workflow id that asked and a hash of the input it sent, the UTC timestamp, the Core Model project with its schema version and the element version in force, the ShEx shape and the mapping it failed, the check that fired and its reason code, and the answer that was refused or warned. Records are retained for the period set in your engagement letter, alongside your other audit artefacts. Your named control owner, and anyone they authorize, can export the pack as JSON at any time.
What is deterministic is the decision, not the prose: for a given contract version and mapping, the same input produces the same refuse or the same warn every time. Your evals assert that decision, not the prose. A generated answer is not reproducible word for word, which is why the boundary, not the model, is what you audit. The engagement maps each boundary check to a control in the framework you already report against — BCBS 239, Solvency II, IFRS 17 and 21 CFR Part 11 are the shapes it takes most often.
When an agent answer becomes a finding, that pack — the refuse or warn record, the approved mapping and the contract version in force — is what your auditors receive, and because the decision is deterministic they can re-run it without us in the room.