Skip to content
Press / to search

Adopt a Harness Contract

Apply ContextOS to one workflow in your existing runtime: define accepted outcomes, bound authority, verify effects, and evaluate a release.

Implementation GuideLast reviewed: Edit on GitHub
At a glance
ContextOS ContractPolicy & GatesCompiler & MemoryAdapter MeshContextActions

Start with one valuable workflow in the runtime you already use. Write down what counts as success, what the agent may change, and how an operator will verify the result. ContextOS supplies the contracts for making those answers inspectable.

This guide is an adoption profile. The published schemas and API Contracts define the versioned interfaces. The repository contains a deterministic context compiler and its tests; you supply production execution, identity, storage, tools, and evaluation infrastructure.

Choose the boundary you need

Your starting pointFirst useful changeEvidence before expanding
A deterministic workflow already meets the taskKeep it; use typed outcomes and evidence where usefulA measured failure or capability gap justifies model reasoning
An agent prototype with toolsWrap one read and one write path with identity, policy, and tool envelopesForbidden actions are denied and claimed success agrees with observed state
A managed agent serviceMap provider hooks and retained session events to your application contractsYou can enforce the required controls and export sufficient evidence
Several runtimes serving the same intentCompare the same task and denial cases across implementationsEquivalent deterministic artifacts and acceptable outcome differences

The five planes assign responsibilities. They do not require five services, a graph database, a memory platform, or a fleet of specialist agents on day one. Planner, Executor, and Critic name responsibilities in the canonical loop; process isolation and additional model calls are deployment choices.

1. Define accepted work

For a support.refund workflow, separate three questions:

  1. Eligibility: does current evidence establish that the order qualifies?
  2. Authorization: may this caller request this refund for this tenant, order, and amount?
  3. Completion: did the payment system actually record the permitted refund?

An eligible refund can still be unauthorized. An authorized call can still time out. A fluent response saying “refunded” establishes neither permission nor completion.

Register the decision’s allowed outcomes and required evidence in a DecisionSpec. Keep operational acceptance cases beside the implementing workflow:

CaseExpected observation
Eligible order, authorized caller, successful provider responseOne permitted effect, confirmed postcondition, evidence-linked DecisionRecord
Caller cannot act on the orderDenial before payment execution; a retained reason
Required evidence is missing or contradictoryNo refund commitment; typed rejection, replan, or escalation
Approval expires before executionAuthorization is rechecked; the expired grant cannot authorize the effect
Provider commits but the response is lostReconcile by the recorded operation/idempotency contract before retrying
Run is cancelled or resumedNo renewed authority, reset budget, or duplicate completed effect

These are example cases, not a new refund policy or schema. Your business owns the eligibility rules and acceptance thresholds.

2. Record the run and its context

Establish a RunContext before execution: caller and agent identity, tenant, trace, limits, and applicable compatibility safety ceiling. Use scoped credentials that the model cannot widen. Bind the run to a versioned Context Pack and the actual tool/policy configuration.

Compile the required order, ownership, and policy evidence into CompiledContext. Retain provenance, admitted and rejected evidence, conflicts, omissions, and budget reports. Retrieved text, tool descriptions, and recalled memory are data; they cannot grant themselves permission.

Use the reference compiler’s fixtures to understand deterministic admission and evidence gates. Production retrieval must provide trustworthy source metadata, and the implementing decision/action boundaries must honor commit_allowed. A schema-valid context by itself does not prove that its source is true.

3. Enforce before the effect

Route the attempted write through the Tool Gateway contract. Evaluate actual arguments against caller authority, tenant scope, policy, current approval, and the capability’s native ActionRisk. The v1 approval-mode label is a compatibility projection, not the authorization verdict.

The enforcement path needs to be unavoidable. An agent with unrestricted payment credentials can bypass a gateway even when its instructions say otherwise. Place credentials, isolation, and effect authorization where the application or provider can enforce them.

After execution, verify the postcondition against the payment system. Preserve the tool result and its evidence refs, trace, idempotency identity, and applicable recovery metadata. If the result is unresolved, record uncertainty and escalate or reconcile; do not manufacture a completed outcome.

4. Accept and retain the outcome

The DecisionRecord ties the decision to evidence, applied policy, approvals, tool lineage, and replay references. Validate its structure and separately check the evidence supporting its outcome.

Retain a ReplayPacket with pinned inputs and recorded tool results. Audit replay substitutes transcripts and disables live effects. Compare deterministic compiler and policy artifacts exactly; evaluate model-dependent stages against declared outcomes and properties. Replay does not reverse a payment or guarantee identical generated prose.

Where a provider cannot expose a model or harness revision, record the available configuration and observation date and state the limitation. Do not report an immutable release pin that the provider did not supply.

5. Evaluate the release

Freeze a baseline, candidate configuration, task set, graders, resource limits, and environment reset procedure. Include the denial, ambiguity, interruption, and lost-response cases above. Evaluate useful authorized completion alongside prohibited effects, unresolved outcomes, human effort, latency, and full cost.

Use repeated trials for model-dependent stages. Keep tuning evidence separate from protected final evaluation. A fixed-model experiment can attribute differences to declared harness changes; a simultaneous model and harness change is a whole-system comparison. Follow the comparison protocol.

Release a bounded cohort only after its acceptance criteria pass, with an owner, kill switch, and achievable fallback. Configuration rollback restores a configuration; recovery from an external effect requires the capability’s own reconciliation or compensation path.

Assign ownership when execution is managed

Record an actual owner for each row. “The provider handles it” requires an exposed interface, documented behavior, and a tested case.

ResponsibilityWhat to demonstrate
Session and sandboxState survives interruption under the provider’s documented recovery behavior
Identity and credentialsA tool resolves credentials for the right caller and tenant; revocation is observed
Context and skillsYou know which controlled artifacts were available and which evidence reached execution
AuthorizationRequired policy and approval checks happen before each reachable effect
Outcome acceptanceApplication state establishes completion independently of the agent’s report
Evidence and operationsRequired events can be exported, retained, redacted, and joined during an incident
Release and fallbackProvider updates are evaluated; an unavailable prior implementation is not claimed as a rollback option

If a required pre-effect hook or evidence surface is unavailable, constrain the workflow, move the effect to an application-owned tool, or select another execution profile. State the narrower claim. A managed session’s durability does not establish end-to-end workflow conformance.

Add complexity from observed failures

Observed failureCandidate changeEvidence it helped
Required evidence is omittedRetrieval/admission rule or context budget changeCorrect coverage with bounded tokens and no new leakage
Multi-session work loses progressDurable checkpoint/handoff profileResume preserves pins, consumed budget, authority, and effect identity
Repeated valid correction is forgottenScoped promotion-aware memoryUseful recall without stale, conflicting, or unauthorized facts
Independent tasks block each otherBounded delegation profileLower completion time without authority expansion or hidden cost
Output passes self-review but fails usersIndependent grader and calibrated rubricBetter observed outcomes on fresh cases
A newer model handles an old failure itselfRemove a compensating prompt, reset, or planning stepQuality and controls hold with less latency or cost

Every candidate goes through the Improvement Loop. Search may propose changes to the harness; release authority remains separate. Deterministic authorization and evidence requirements cannot be traded away for a better aggregate score.

Evidence behind this approach

This adoption sequence is ContextOS’s synthesis, not a vendor compatibility certification.

Continue with the worked example

Use the Quickstart to model the refund contracts in detail, How It Works to follow the full run, and Harness Engineering to understand the cross-plane discipline. Use the harness audit to assess evidence in an existing implementation.