Adopt a Harness Contract
Apply ContextOS to one workflow in your existing runtime: define accepted outcomes, bound authority, verify effects, and evaluate a release.
Start with one valuable workflow in the runtime you already use. Write down what counts as success, what the agent may change, and how an operator will verify the result. ContextOS supplies the contracts for making those answers inspectable.
This guide is an adoption profile. The published schemas and API Contracts define the versioned interfaces. The repository contains a deterministic context compiler and its tests; you supply production execution, identity, storage, tools, and evaluation infrastructure.
Choose the boundary you need
| Your starting point | First useful change | Evidence before expanding |
|---|---|---|
| A deterministic workflow already meets the task | Keep it; use typed outcomes and evidence where useful | A measured failure or capability gap justifies model reasoning |
| An agent prototype with tools | Wrap one read and one write path with identity, policy, and tool envelopes | Forbidden actions are denied and claimed success agrees with observed state |
| A managed agent service | Map provider hooks and retained session events to your application contracts | You can enforce the required controls and export sufficient evidence |
| Several runtimes serving the same intent | Compare the same task and denial cases across implementations | Equivalent deterministic artifacts and acceptable outcome differences |
The five planes assign responsibilities. They do not require five services, a graph database, a memory platform, or a fleet of specialist agents on day one. Planner, Executor, and Critic name responsibilities in the canonical loop; process isolation and additional model calls are deployment choices.
1. Define accepted work
For a support.refund workflow, separate three questions:
- Eligibility: does current evidence establish that the order qualifies?
- Authorization: may this caller request this refund for this tenant, order, and amount?
- Completion: did the payment system actually record the permitted refund?
An eligible refund can still be unauthorized. An authorized call can still time out. A fluent response saying “refunded” establishes neither permission nor completion.
Register the decision’s allowed outcomes and required evidence in a DecisionSpec. Keep operational acceptance cases beside the implementing workflow:
| Case | Expected observation |
|---|---|
| Eligible order, authorized caller, successful provider response | One permitted effect, confirmed postcondition, evidence-linked DecisionRecord |
| Caller cannot act on the order | Denial before payment execution; a retained reason |
| Required evidence is missing or contradictory | No refund commitment; typed rejection, replan, or escalation |
| Approval expires before execution | Authorization is rechecked; the expired grant cannot authorize the effect |
| Provider commits but the response is lost | Reconcile by the recorded operation/idempotency contract before retrying |
| Run is cancelled or resumed | No renewed authority, reset budget, or duplicate completed effect |
These are example cases, not a new refund policy or schema. Your business owns the eligibility rules and acceptance thresholds.
2. Record the run and its context
Establish a RunContext before execution: caller and agent identity, tenant, trace, limits, and applicable compatibility safety ceiling. Use scoped credentials that the model cannot widen. Bind the run to a versioned Context Pack and the actual tool/policy configuration.
Compile the required order, ownership, and policy evidence into CompiledContext. Retain provenance, admitted and rejected evidence, conflicts, omissions, and budget reports. Retrieved text, tool descriptions, and recalled memory are data; they cannot grant themselves permission.
Use the reference compiler’s fixtures to understand deterministic admission and evidence gates. Production retrieval must provide trustworthy source metadata, and the implementing decision/action boundaries must honor commit_allowed. A schema-valid context by itself does not prove that its source is true.
3. Enforce before the effect
Route the attempted write through the Tool Gateway contract. Evaluate actual arguments against caller authority, tenant scope, policy, current approval, and the capability’s native ActionRisk. The v1 approval-mode label is a compatibility projection, not the authorization verdict.
The enforcement path needs to be unavoidable. An agent with unrestricted payment credentials can bypass a gateway even when its instructions say otherwise. Place credentials, isolation, and effect authorization where the application or provider can enforce them.
After execution, verify the postcondition against the payment system. Preserve the tool result and its evidence refs, trace, idempotency identity, and applicable recovery metadata. If the result is unresolved, record uncertainty and escalate or reconcile; do not manufacture a completed outcome.
4. Accept and retain the outcome
The DecisionRecord ties the decision to evidence, applied policy, approvals, tool lineage, and replay references. Validate its structure and separately check the evidence supporting its outcome.
Retain a ReplayPacket with pinned inputs and recorded tool results. Audit replay substitutes transcripts and disables live effects. Compare deterministic compiler and policy artifacts exactly; evaluate model-dependent stages against declared outcomes and properties. Replay does not reverse a payment or guarantee identical generated prose.
Where a provider cannot expose a model or harness revision, record the available configuration and observation date and state the limitation. Do not report an immutable release pin that the provider did not supply.
5. Evaluate the release
Freeze a baseline, candidate configuration, task set, graders, resource limits, and environment reset procedure. Include the denial, ambiguity, interruption, and lost-response cases above. Evaluate useful authorized completion alongside prohibited effects, unresolved outcomes, human effort, latency, and full cost.
Use repeated trials for model-dependent stages. Keep tuning evidence separate from protected final evaluation. A fixed-model experiment can attribute differences to declared harness changes; a simultaneous model and harness change is a whole-system comparison. Follow the comparison protocol.
Release a bounded cohort only after its acceptance criteria pass, with an owner, kill switch, and achievable fallback. Configuration rollback restores a configuration; recovery from an external effect requires the capability’s own reconciliation or compensation path.
Assign ownership when execution is managed
Record an actual owner for each row. “The provider handles it” requires an exposed interface, documented behavior, and a tested case.
| Responsibility | What to demonstrate |
|---|---|
| Session and sandbox | State survives interruption under the provider’s documented recovery behavior |
| Identity and credentials | A tool resolves credentials for the right caller and tenant; revocation is observed |
| Context and skills | You know which controlled artifacts were available and which evidence reached execution |
| Authorization | Required policy and approval checks happen before each reachable effect |
| Outcome acceptance | Application state establishes completion independently of the agent’s report |
| Evidence and operations | Required events can be exported, retained, redacted, and joined during an incident |
| Release and fallback | Provider updates are evaluated; an unavailable prior implementation is not claimed as a rollback option |
If a required pre-effect hook or evidence surface is unavailable, constrain the workflow, move the effect to an application-owned tool, or select another execution profile. State the narrower claim. A managed session’s durability does not establish end-to-end workflow conformance.
Add complexity from observed failures
| Observed failure | Candidate change | Evidence it helped |
|---|---|---|
| Required evidence is omitted | Retrieval/admission rule or context budget change | Correct coverage with bounded tokens and no new leakage |
| Multi-session work loses progress | Durable checkpoint/handoff profile | Resume preserves pins, consumed budget, authority, and effect identity |
| Repeated valid correction is forgotten | Scoped promotion-aware memory | Useful recall without stale, conflicting, or unauthorized facts |
| Independent tasks block each other | Bounded delegation profile | Lower completion time without authority expansion or hidden cost |
| Output passes self-review but fails users | Independent grader and calibrated rubric | Better observed outcomes on fresh cases |
| A newer model handles an old failure itself | Remove a compensating prompt, reset, or planning step | Quality and controls hold with less latency or cost |
Every candidate goes through the Improvement Loop. Search may propose changes to the harness; release authority remains separate. Deterministic authorization and evidence requirements cannot be traded away for a better aggregate score.
Evidence behind this approach
This adoption sequence is ContextOS’s synthesis, not a vendor compatibility certification.
- OpenAI’s harness engineering report emphasizes legible environments, repository knowledge, and mechanical feedback. Its coding experience informs environment design, not a universal productivity forecast.
- Anthropic’s long-running harness report motivates durable progress artifacts and verification across sessions.
- Anthropic’s managed-agent design separates session, harness, and sandbox and explains why compensating scaffolding can become unnecessary as models improve.
- Anthropic’s agent evaluation guidance distinguishes transcripts from environment outcomes and recommends task-appropriate graders.
- LangChain’s Deep Agents v0.7 report illustrates measured harness simplification and differing effects across models.
- Meta-Harness research investigates harness search. Benchmark improvements motivate controlled experiments; they do not establish safety or transfer to your workflow.
Continue with the worked example
Use the Quickstart to model the refund contracts in detail, How It Works to follow the full run, and Harness Engineering to understand the cross-plane discipline. Use the harness audit to assess evidence in an existing implementation.