ContextOS 2026
ContextOS's direction as agent harnesses evolve: portable control contracts, measured simplification, managed-runtime ownership, and evidence-gated conformance milestones.
ContextOS’s public pitch is an open contract for reliable agent harnesses. The portable decision-control, evidence, and conformance specification is its technical foundation. Keep the five-plane architecture and v1 contracts while making adoption useful inside existing runtimes: define accepted work, preserve authority across execution boundaries, verify effects, and retain evidence for release decisions.
Direction reviewed on October 3, 2026
Three developments shape the approach. Managed harnesses separate session, execution, and sandbox interfaces. Harness simplification is being evaluated across models. Harness search is a research target. These signals support stable control contracts around changing implementations; they do not demonstrate ContextOS interoperability or universal performance gains.
The practical changes are:
- Outcome first: adopt one workflow from an observable acceptance contract, with success, denial, ambiguity, interruption, and recovery cases.
- Own the seams: identify which provider or application enforces context, identity, authorization, outcome acceptance, retention, and fallback. Unavailable hooks limit the workflow’s conformance claim.
- Re-evaluate complexity: compare the model, harness, and environment. Remove compensating prompts or orchestration when they no longer help; preserve deterministic control obligations.
- Publish implementation evidence: prioritize executable fixtures and transparent comparison results over more architecture vocabulary. A schema or blueprint is not an end-to-end production result.
The adoption guide turns this direction into a workflow sequence. No new runtime type or breaking envelope is introduced by this positioning review.
What changed since the original framework
| Earlier assumption | 2026 reality | ContextOS revision |
|---|---|---|
| One model is the agent | Models are routed by capability, risk, cost, residency, and measured quality | Model profiles and deterministic routing lineage; the documented RoutingDecision shape remains a candidate profile |
| Tools are API functions | Agents also use browsers, computers, agent peers, and asynchronous tasks | Interaction boundary is part of action risk and audit |
| Risk fits one tier | Effect, authority, reversibility, interaction, and data exposure vary independently | Multidimensional ActionRisk; ApprovalMode becomes a compatibility projection |
| A run ends with one response | Work pauses, checkpoints, waits for approvals, and resumes | Durable session state with pinned artifacts and monotonic budgets |
| More context is better | Long context can amplify stale, hostile, irrelevant, or mutually contradictory content | Proof-carrying materialized views with provenance, omissions, conflicts, and sufficiency gates |
| Multi-agent means parallel prompts | Delegation creates a new authority and budget boundary | Parent-issued child claims, scoped budgets, isolated effects, merge verdicts |
| Evaluation happens after launch | Harnesses can search prompts, routing, retrieval, and orchestration continuously | Proposal-based improvement with holdouts, release gates, and rollback |
The seven 2026 invariants
- Authority is explicit. Every model call, tool call, delegation, and approval names the principal chain and scope ceiling.
- Risk is multidimensional. Policy evaluates effect, authority, reversibility, interaction, and data classification independently.
- Context is a proof-carrying materialized view. Every material block has provenance, trust treatment, freshness, and a reason for inclusion; omissions and conflicts remain visible, and commit depends on explicit sufficiency.
- Execution is durable. A long-running run can checkpoint and resume without repeating effects or silently refreshing pinned inputs.
- Delegation narrows. A child agent receives a strict subset of the parent’s authority, tools, data scope, time, and budget.
- Models are replaceable dependencies. Routing is deterministic for identical recorded inputs, and every fallback preserves the requested contract.
- Change is governed. Agents may propose improvements; they do not silently promote policy, memory, prompts, tools, evaluators, or routing rules.
Revised contract stack
RunContext
authority + tenant + trace + monotonic budget + session pins
-> ContextPack
policy + evidence requirements + tools + memory + decision specs
-> CompiledContext
selected blocks + provenance + omissions + conflicts + sufficiency + controls + hash
-> documented RoutingDecision profile + bounded decision loop
model profile + plan + critic verdicts + delegation lanes
-> ToolEnvelope / AgentDelegationEnvelope
ActionRisk + principal chain + policy decision + idempotency
-> DecisionRecord + ReplayPacket
outcome + effects + evidence + lineage + scorecard + replay stateThe reference compiler remains deterministic. Model inference may be probabilistic, but artifact selection, policy evaluation, budget checks, capability filtering, and threshold comparisons are reproducible from recorded inputs.
Reference implementation status
The framework intentionally distinguishes a published contract from production integration:
| 2026 capability | Repository status | Boundary |
|---|---|---|
| Context provenance, integrity, freshness, instruction treatment | Typed and emitted by the reference compiler; represented in the public CompiledContext schema | Production retrieval must supply structured candidates and trustworthy source metadata |
| Context admission, sufficiency-aware evidence gates, typed conflicts, structured omissions | Typed, schema-published, compiler-enforced, and covered by conformance tests | Production retrieval must emit conflict markers; Decision and Action runtimes must honor commit_allowed |
Native ActionRisk | Typed and represented in Context Pack and Tool Envelope schemas | Just-in-time enforcement against actual tool arguments belongs at the production Tool Gateway |
AgentDelegationEnvelope | Typed and schema-published | Durable child execution, cancellation, and effect isolation require an orchestrator implementation |
ReplayPacket | Typed and schema-published | Replay runners and transcript stores remain deployment concerns |
| Durable session checkpoints | Documented execution profile; no canonical SessionCheckpoint schema or exported type exists | A production orchestrator must prove resume, cancellation, expiry, pin preservation, effect deduplication, and monotonic budgets before promotion |
Model profiles and RoutingDecision | Documented in the AI Gateway and LLM Router and control-plane docs; no canonical exported RoutingDecision type exists | A production gateway must persist routing lineage; promotion requires a validated schema and conformance fixtures |
| Complete release tuple and proof-carrying execution packet | Documented aggregate protocol in Harness Engineering | Composed from canonical pins, envelopes, records, traces, and receipts; not a new core interface |
| Observable trajectories, protocol cards, and paired Skill Lift | Documented evaluation protocol in Evaluation and Observability | Evaluation stores observable runtime evidence, not hidden chain-of-thought; no canonical Trajectory type yet |
| Insight, strategy, feedback, note, research, and tuning records | Candidate improvement profile; the six artifact names are not canonical exported types | Implementations version and validate local shapes; promotion requires independent evaluation, authority, rollback, and a future schema decision |
| Lifecycle security and effective-authority graph | Documented release gate and operational control view | Production CI/CD, identity systems, gateways, and revocation paths must enforce them |
“Typed” does not mean “operationally enforced everywhere.” Use the compiler and public schemas as conformance references, then close the production boundaries named in the last column.
Strategic choice for late 2026
Hosted runtimes increasingly provide model calls, tool execution, sessions, browsers, sandboxes, and orchestration. ContextOS should not duplicate that commodity plumbing. Its durable value is the contract at the seams:
- compile proof-carrying context before a model call;
- compute effective authority before a model, tool, or agent handoff;
- evaluate multidimensional action risk immediately before an effect;
- bind protocol activity to portable evidence, outcome, and replay records; and
- test whether different implementations preserve those guarantees.
This makes ContextOS useful across vendor SDKs, managed agent platforms, MCP servers, A2A peers, and local harnesses. Interoperability protocols move messages; ContextOS proves whether the resulting decision was governed.
A September 1 research signal reinforces the direction without changing the contract: the Harness-of-Harness preprint reports gains from small verifiable increments, independent evaluation, versioned histories, and output constraints across multi-iteration software work. ContextOS treats those findings as evidence for the conformance and improvement workstreams, not as proof that autonomous self-modification or a new core type is safe.
Standards alignment snapshot
ContextOS should profile and compose current standards instead of inventing competing transports or telemetry:
| External surface | Late-2026 state | ContextOS position |
|---|---|---|
| MCP 2026-07-28 | Stateless requests, per-request negotiation, multi-round-trip input, optional Tasks extension; Roots/Sampling/Logging deprecated | Publish a current adapter profile and isolate the 2025 lifecycle as legacy compatibility |
| A2A 1.0 | Agent discovery, messages, artifacts, durable task lifecycle, authentication, versioning | Map each hop to a narrowing AgentDelegationEnvelope; an Agent Card is discovery, not trust |
| OWASP Agent Control Standard | Runtime hooks and interoperable control/telemetry vocabulary in public preview | Integrate ContextOS policy and evidence at the hook boundary; do not create a rival hook protocol |
| OpenTelemetry GenAI conventions | Shared semantic attributes continue to evolve | Keep portable W3C trace context and version any ContextOS-specific attributes |
| NIST AI Agent Standards Initiative | Focus on agent security, identity, interoperability, and evaluation | Contribute conformance artifacts and map controls; avoid claiming certification |
| SLSA, Sigstore, and in-toto | Mature software-supply-chain building blocks | Reuse attestations for packs, adapters, policies, and runtime artifacts rather than creating a proprietary signing ecosystem |
The standards snapshot is informative and date-sensitive. The linked standards remain authoritative.
Evidence-gated roadmap
The earlier September–December windows were planning targets, not delivery claims. The milestones below replace calendar promises. The reference-status table above describes available artifacts; cross-runtime conformance remains an exit criterion to demonstrate, not an achieved result asserted by this page.
| Milestone | Outcome | Deliverables | Exit evidence |
|---|---|---|---|
| 1. Coherent adoption | Clear positioning and contract boundaries | Outcome-first reading path, status model, ActionRisk/v1 language, link and terminology checks | The worked workflow distinguishes reference artifacts from production enforcement; content and contract tests pass |
| 2. Portable conformance | Reproducible comparison across implementations | Golden packs, compile fixtures, envelope/denial cases, replay inputs, explicit provider limitations | At least two independent harnesses produce equivalent deterministic artifacts for the same fixtures; publish failures and test conditions |
| 3. Operational acceptance | Validated identity, delegation, and effect profiles | Identity and A2A mappings, postcondition/recovery cases, control/telemetry crosswalk | Reviewed traces demonstrate non-escalation, revocation, idempotency, cancellation, and partial-effect handling |
| 4. Version decision | Evidence-supported contract evolution | Compatibility telemetry, migration guide, rejected alternatives, validated schema proposals | Promote only changes with implementation evidence, fixtures, and bounded migration cost; retain v1 where evidence is insufficient |
Workstreams
- Spec semantics: remove ambiguous total-order risk language and define the authority/effect/evidence invariants once.
- Conformance: turn prose obligations into fixtures, negative cases, and deterministic comparison rules.
- Protocol profiles: maintain explicit MCP-current, MCP-legacy, and A2A mappings without importing protocol state into the core contract.
- Identity and delegation: connect principal chains to deployable OAuth/workload-identity patterns while keeping credentials out of model context.
- Effect evidence: require observable postconditions, idempotency, and recovery metadata using existing envelopes before proposing a new receipt type.
- Operational proof: publish replay and cross-runtime results, not feature-count claims.
Explicit non-goals
- Building another full agent orchestration SDK or managed runtime.
- Creating proprietary replacements for MCP, A2A, OpenTelemetry, OAuth, or supply-chain attestations.
- Storing hidden chain-of-thought as an audit requirement; observable actions, inputs, outputs, verdicts, and state diffs are the evidence surface.
- Turning
ActionRiskinto a single numeric score. - Adding
EffectReceipt,OutcomeVerification,Trajectory, orRoutingDecisionto the canonical type surface before validation. - Publishing a breaking envelope merely to satisfy a roadmap date.
Context sufficiency and conflict state
The earlier framework used evidence coverage as the precondition for a governed decision. Coverage is necessary, but sufficient-context research shows that relevance alone does not establish completeness and that contradictory context is insufficient. A context can contain a relevant passage for every requested field and still be incomplete, inconclusive, or contradictory. ACL 2026 temporal-conflict results further show that models may describe a temporal conflict correctly and then fail to apply it to the final prediction.
ContextOS therefore treats CompiledContext as a proof-carrying materialized view over the pinned pack, request, evidence, memory, and policy state:
ContextProvenanceexplains where each admitted block came from and how it may be used;ContextOmissionpreserves what budget pressure excluded;EvidenceConflictMarkerpreserves competing evidence branches and the requirements they affect;EvidenceGate.sufficiencysummarizes coverage, admission, and conflict state; andcommit_allowedis derived from that state, not from model confidence.
An unresolved required conflict fails closed. A resolved conflict must name the resolution artifact and selected evidence refs; the losing branch remains auditable as superseded_by_resolution. This implements the hard-fail behavior already required by the Knowledge Graph instead of asking the model to improvise source precedence.
This is a bounded first step toward the claim-level execution provenance identified by the 2026 evidence tracing and execution provenance survey. A complete semantic graph connecting retrieved evidence, intermediate claims, actions, and final outputs remains a candidate extension until its schema and evaluation protocol are validated. The current contract closes the concrete spec gap without inventing an untested sixth plane.
Native action risk
The original ApprovalMode values—read_only, local_write, network, delegated, and destructive—are useful operational labels but not a coherent ordering. network describes location, delegated describes authority, and destructive describes reversibility.
The 2026 contract adds ActionRisk:
{
"effect": "external_state",
"authority": "user_delegated",
"reversibility": "compensatable",
"interaction": "browser",
"data_scope": "CONFIDENTIAL",
"decision_ttl_seconds": 120
}Policy checks each dimension and the actual arguments immediately before execution. The legacy approval mode remains in the v1 ToolCallEnvelope; native risk fields are additive in v1. Making them required in a future major envelope is an evidence-gated migration decision, not a promised release.
Durable and delegated execution
A durable session pins its Context Pack, policy bundles, knowledge snapshot, model-routing policy, evaluator suite, and tool registry view. Resume restores atomic budget usage and continues after the last committed Critic verdict. It never replays a completed effect merely because a process restarted.
A delegation is a derived run, not an informal message. The parent issues a child identity claim with:
- allowed intents and output schema;
- capability allow-list and argument ceilings;
- data classification ceiling;
- token, cost, tool-call, and wall-clock sub-budgets;
- expiration and cancellation semantics;
- a rule that child effects remain isolated until the parent Critic accepts them.
The parent’s authority is an upper bound. Delegation cannot mint permission, reset consumed budget, or hide model/tool calls from the parent trace.
Multimodal and computer-use agents
Screenshots, audio, documents, and UI state enter the Context plane as content-addressed artifacts. Extracted text is derived evidence and retains a pointer to the source artifact and extraction model. Instructions found inside retrieved content or a UI are untrusted data unless policy explicitly promotes them.
Computer-use actions require a pre-action observation, a bounded action proposal, a post-action observation, and an environment receipt. Sensitive inputs are referenced through secret handles, not exposed to the model as plaintext. For irreversible or ambiguous UI actions, the Critic must prefer a human handoff over coordinate guessing.
Adoption path
| Stage | Required change | Compatibility |
|---|---|---|
| 0. Inventory | List model calls, tools, background jobs, agent peers, browser/computer surfaces, and current approvals | No runtime change |
| 1. Observe | Emit principal chains, routing lineage, context ledgers, conflict markers, tool-result postcondition/recovery evidence, and replay packets | Additive telemetry |
| 2. Classify | Add ActionRisk to capabilities, validate conservative legacy-mode projections, and map conflicts to evidence requirements | v1 envelopes remain accepted |
| 3. Enforce | Apply per-dimension policy and just-in-time argument checks | Mixed v1/native adapters supported |
| 4. Durabilize | Add checkpoints, artifact pins, cancellation, and monotonic budget restoration | Synchronous runs unchanged |
| 5. Delegate | Introduce child claims, sub-budgets, isolated effects, and merge verdicts | No implicit agent-to-agent authority |
| 6. Promote | Require replay, holdouts, release gates, and rollback for harness changes | Manual promotion remains valid |
What is deliberately not changed
- The five planes remain the architecture map; new capabilities fit within them.
- Policy stays outside model reasoning.
- The Tool Gateway remains the only path to external effects.
- Durable memory still requires promotion and provenance.
DecisionRecordremains the audit receipt, and replay remains the test of operational truth.- Protocols such as MCP and A2A describe interoperability; they do not grant authority.
Review checklist
A 2026-ready workflow can answer all of these with typed artifacts:
- Which principal and delegated authority permitted each effect?
- What were the action’s effect, reversibility, interaction, and data boundaries?
- What exact context did each model see, and which candidate context was omitted?
- Which required claims were conflicted, how were they resolved, and which evidence branch was superseded?
- Which model profile was selected, why, and which fallbacks were rejected?
- Can the run resume without refreshing pins or repeating side effects?
- Did every child agent receive narrower authority and a real sub-budget?
- Can the decision be replayed with side effects disabled?
- Can every promoted harness change be rolled back independently?
Success measures
The late-2026 program succeeds when:
- two or more runtimes pass the same deterministic conformance fixtures;
- every externally visible effect is linked to policy, identity, idempotency, postcondition, and recovery evidence;
- native-risk coverage rises without breaking v1 envelope consumers;
- protocol upgrades are isolated profile changes rather than core-contract rewrites;
- replay can compare outcomes without re-executing live effects; and
- proposed core types are rejected as readily as they are added when implementation evidence is weak.
Next steps
- Use Foundations for the plane-by-plane operating model.
- Use Governance for native action risk and compatibility approval modes.
- Use Orchestration for durable sessions and subagent lanes.
- Use AI Gateway and LLM Router for governed model selection.
- Use API Contracts and Runtime API Schemas for executable envelopes.
- Use Research Synthesis for the frozen 109-essay evidence-to-spec snapshot through 2026-08-28 and the boundary between canonical contracts, documented protocols, and candidate extensions.