Skip to content
Press / to search

ContextOS 2026

ContextOS's direction as agent harnesses evolve: portable control contracts, measured simplification, managed-runtime ownership, and evidence-gated conformance milestones.

Framework RevisionLast reviewed: Edit on GitHub
At a glance
ContextOS ContractPolicy & GatesCompiler & MemoryAdapter MeshContextActions

ContextOS’s public pitch is an open contract for reliable agent harnesses. The portable decision-control, evidence, and conformance specification is its technical foundation. Keep the five-plane architecture and v1 contracts while making adoption useful inside existing runtimes: define accepted work, preserve authority across execution boundaries, verify effects, and retain evidence for release decisions.

Direction reviewed on October 3, 2026

Three developments shape the approach. Managed harnesses separate session, execution, and sandbox interfaces. Harness simplification is being evaluated across models. Harness search is a research target. These signals support stable control contracts around changing implementations; they do not demonstrate ContextOS interoperability or universal performance gains.

The practical changes are:

  • Outcome first: adopt one workflow from an observable acceptance contract, with success, denial, ambiguity, interruption, and recovery cases.
  • Own the seams: identify which provider or application enforces context, identity, authorization, outcome acceptance, retention, and fallback. Unavailable hooks limit the workflow’s conformance claim.
  • Re-evaluate complexity: compare the model, harness, and environment. Remove compensating prompts or orchestration when they no longer help; preserve deterministic control obligations.
  • Publish implementation evidence: prioritize executable fixtures and transparent comparison results over more architecture vocabulary. A schema or blueprint is not an end-to-end production result.

The adoption guide turns this direction into a workflow sequence. No new runtime type or breaking envelope is introduced by this positioning review.

What changed since the original framework

Earlier assumption2026 realityContextOS revision
One model is the agentModels are routed by capability, risk, cost, residency, and measured qualityModel profiles and deterministic routing lineage; the documented RoutingDecision shape remains a candidate profile
Tools are API functionsAgents also use browsers, computers, agent peers, and asynchronous tasksInteraction boundary is part of action risk and audit
Risk fits one tierEffect, authority, reversibility, interaction, and data exposure vary independentlyMultidimensional ActionRisk; ApprovalMode becomes a compatibility projection
A run ends with one responseWork pauses, checkpoints, waits for approvals, and resumesDurable session state with pinned artifacts and monotonic budgets
More context is betterLong context can amplify stale, hostile, irrelevant, or mutually contradictory contentProof-carrying materialized views with provenance, omissions, conflicts, and sufficiency gates
Multi-agent means parallel promptsDelegation creates a new authority and budget boundaryParent-issued child claims, scoped budgets, isolated effects, merge verdicts
Evaluation happens after launchHarnesses can search prompts, routing, retrieval, and orchestration continuouslyProposal-based improvement with holdouts, release gates, and rollback

The seven 2026 invariants

  1. Authority is explicit. Every model call, tool call, delegation, and approval names the principal chain and scope ceiling.
  2. Risk is multidimensional. Policy evaluates effect, authority, reversibility, interaction, and data classification independently.
  3. Context is a proof-carrying materialized view. Every material block has provenance, trust treatment, freshness, and a reason for inclusion; omissions and conflicts remain visible, and commit depends on explicit sufficiency.
  4. Execution is durable. A long-running run can checkpoint and resume without repeating effects or silently refreshing pinned inputs.
  5. Delegation narrows. A child agent receives a strict subset of the parent’s authority, tools, data scope, time, and budget.
  6. Models are replaceable dependencies. Routing is deterministic for identical recorded inputs, and every fallback preserves the requested contract.
  7. Change is governed. Agents may propose improvements; they do not silently promote policy, memory, prompts, tools, evaluators, or routing rules.

Revised contract stack

RunContext
  authority + tenant + trace + monotonic budget + session pins
    -> ContextPack
       policy + evidence requirements + tools + memory + decision specs
    -> CompiledContext
       selected blocks + provenance + omissions + conflicts + sufficiency + controls + hash
    -> documented RoutingDecision profile + bounded decision loop
       model profile + plan + critic verdicts + delegation lanes
    -> ToolEnvelope / AgentDelegationEnvelope
       ActionRisk + principal chain + policy decision + idempotency
    -> DecisionRecord + ReplayPacket
       outcome + effects + evidence + lineage + scorecard + replay state

The reference compiler remains deterministic. Model inference may be probabilistic, but artifact selection, policy evaluation, budget checks, capability filtering, and threshold comparisons are reproducible from recorded inputs.

Reference implementation status

The framework intentionally distinguishes a published contract from production integration:

2026 capabilityRepository statusBoundary
Context provenance, integrity, freshness, instruction treatmentTyped and emitted by the reference compiler; represented in the public CompiledContext schemaProduction retrieval must supply structured candidates and trustworthy source metadata
Context admission, sufficiency-aware evidence gates, typed conflicts, structured omissionsTyped, schema-published, compiler-enforced, and covered by conformance testsProduction retrieval must emit conflict markers; Decision and Action runtimes must honor commit_allowed
Native ActionRiskTyped and represented in Context Pack and Tool Envelope schemasJust-in-time enforcement against actual tool arguments belongs at the production Tool Gateway
AgentDelegationEnvelopeTyped and schema-publishedDurable child execution, cancellation, and effect isolation require an orchestrator implementation
ReplayPacketTyped and schema-publishedReplay runners and transcript stores remain deployment concerns
Durable session checkpointsDocumented execution profile; no canonical SessionCheckpoint schema or exported type existsA production orchestrator must prove resume, cancellation, expiry, pin preservation, effect deduplication, and monotonic budgets before promotion
Model profiles and RoutingDecisionDocumented in the AI Gateway and LLM Router and control-plane docs; no canonical exported RoutingDecision type existsA production gateway must persist routing lineage; promotion requires a validated schema and conformance fixtures
Complete release tuple and proof-carrying execution packetDocumented aggregate protocol in Harness EngineeringComposed from canonical pins, envelopes, records, traces, and receipts; not a new core interface
Observable trajectories, protocol cards, and paired Skill LiftDocumented evaluation protocol in Evaluation and ObservabilityEvaluation stores observable runtime evidence, not hidden chain-of-thought; no canonical Trajectory type yet
Insight, strategy, feedback, note, research, and tuning recordsCandidate improvement profile; the six artifact names are not canonical exported typesImplementations version and validate local shapes; promotion requires independent evaluation, authority, rollback, and a future schema decision
Lifecycle security and effective-authority graphDocumented release gate and operational control viewProduction CI/CD, identity systems, gateways, and revocation paths must enforce them

“Typed” does not mean “operationally enforced everywhere.” Use the compiler and public schemas as conformance references, then close the production boundaries named in the last column.

Strategic choice for late 2026

Hosted runtimes increasingly provide model calls, tool execution, sessions, browsers, sandboxes, and orchestration. ContextOS should not duplicate that commodity plumbing. Its durable value is the contract at the seams:

  • compile proof-carrying context before a model call;
  • compute effective authority before a model, tool, or agent handoff;
  • evaluate multidimensional action risk immediately before an effect;
  • bind protocol activity to portable evidence, outcome, and replay records; and
  • test whether different implementations preserve those guarantees.

This makes ContextOS useful across vendor SDKs, managed agent platforms, MCP servers, A2A peers, and local harnesses. Interoperability protocols move messages; ContextOS proves whether the resulting decision was governed.

A September 1 research signal reinforces the direction without changing the contract: the Harness-of-Harness preprint reports gains from small verifiable increments, independent evaluation, versioned histories, and output constraints across multi-iteration software work. ContextOS treats those findings as evidence for the conformance and improvement workstreams, not as proof that autonomous self-modification or a new core type is safe.

Standards alignment snapshot

ContextOS should profile and compose current standards instead of inventing competing transports or telemetry:

External surfaceLate-2026 stateContextOS position
MCP 2026-07-28Stateless requests, per-request negotiation, multi-round-trip input, optional Tasks extension; Roots/Sampling/Logging deprecatedPublish a current adapter profile and isolate the 2025 lifecycle as legacy compatibility
A2A 1.0Agent discovery, messages, artifacts, durable task lifecycle, authentication, versioningMap each hop to a narrowing AgentDelegationEnvelope; an Agent Card is discovery, not trust
OWASP Agent Control StandardRuntime hooks and interoperable control/telemetry vocabulary in public previewIntegrate ContextOS policy and evidence at the hook boundary; do not create a rival hook protocol
OpenTelemetry GenAI conventionsShared semantic attributes continue to evolveKeep portable W3C trace context and version any ContextOS-specific attributes
NIST AI Agent Standards InitiativeFocus on agent security, identity, interoperability, and evaluationContribute conformance artifacts and map controls; avoid claiming certification
SLSA, Sigstore, and in-totoMature software-supply-chain building blocksReuse attestations for packs, adapters, policies, and runtime artifacts rather than creating a proprietary signing ecosystem

The standards snapshot is informative and date-sensitive. The linked standards remain authoritative.

Evidence-gated roadmap

The earlier September–December windows were planning targets, not delivery claims. The milestones below replace calendar promises. The reference-status table above describes available artifacts; cross-runtime conformance remains an exit criterion to demonstrate, not an achieved result asserted by this page.

MilestoneOutcomeDeliverablesExit evidence
1. Coherent adoptionClear positioning and contract boundariesOutcome-first reading path, status model, ActionRisk/v1 language, link and terminology checksThe worked workflow distinguishes reference artifacts from production enforcement; content and contract tests pass
2. Portable conformanceReproducible comparison across implementationsGolden packs, compile fixtures, envelope/denial cases, replay inputs, explicit provider limitationsAt least two independent harnesses produce equivalent deterministic artifacts for the same fixtures; publish failures and test conditions
3. Operational acceptanceValidated identity, delegation, and effect profilesIdentity and A2A mappings, postcondition/recovery cases, control/telemetry crosswalkReviewed traces demonstrate non-escalation, revocation, idempotency, cancellation, and partial-effect handling
4. Version decisionEvidence-supported contract evolutionCompatibility telemetry, migration guide, rejected alternatives, validated schema proposalsPromote only changes with implementation evidence, fixtures, and bounded migration cost; retain v1 where evidence is insufficient

Workstreams

  1. Spec semantics: remove ambiguous total-order risk language and define the authority/effect/evidence invariants once.
  2. Conformance: turn prose obligations into fixtures, negative cases, and deterministic comparison rules.
  3. Protocol profiles: maintain explicit MCP-current, MCP-legacy, and A2A mappings without importing protocol state into the core contract.
  4. Identity and delegation: connect principal chains to deployable OAuth/workload-identity patterns while keeping credentials out of model context.
  5. Effect evidence: require observable postconditions, idempotency, and recovery metadata using existing envelopes before proposing a new receipt type.
  6. Operational proof: publish replay and cross-runtime results, not feature-count claims.

Explicit non-goals

  • Building another full agent orchestration SDK or managed runtime.
  • Creating proprietary replacements for MCP, A2A, OpenTelemetry, OAuth, or supply-chain attestations.
  • Storing hidden chain-of-thought as an audit requirement; observable actions, inputs, outputs, verdicts, and state diffs are the evidence surface.
  • Turning ActionRisk into a single numeric score.
  • Adding EffectReceipt, OutcomeVerification, Trajectory, or RoutingDecision to the canonical type surface before validation.
  • Publishing a breaking envelope merely to satisfy a roadmap date.

Context sufficiency and conflict state

The earlier framework used evidence coverage as the precondition for a governed decision. Coverage is necessary, but sufficient-context research shows that relevance alone does not establish completeness and that contradictory context is insufficient. A context can contain a relevant passage for every requested field and still be incomplete, inconclusive, or contradictory. ACL 2026 temporal-conflict results further show that models may describe a temporal conflict correctly and then fail to apply it to the final prediction.

ContextOS therefore treats CompiledContext as a proof-carrying materialized view over the pinned pack, request, evidence, memory, and policy state:

  • ContextProvenance explains where each admitted block came from and how it may be used;
  • ContextOmission preserves what budget pressure excluded;
  • EvidenceConflictMarker preserves competing evidence branches and the requirements they affect;
  • EvidenceGate.sufficiency summarizes coverage, admission, and conflict state; and
  • commit_allowed is derived from that state, not from model confidence.

An unresolved required conflict fails closed. A resolved conflict must name the resolution artifact and selected evidence refs; the losing branch remains auditable as superseded_by_resolution. This implements the hard-fail behavior already required by the Knowledge Graph instead of asking the model to improvise source precedence.

This is a bounded first step toward the claim-level execution provenance identified by the 2026 evidence tracing and execution provenance survey. A complete semantic graph connecting retrieved evidence, intermediate claims, actions, and final outputs remains a candidate extension until its schema and evaluation protocol are validated. The current contract closes the concrete spec gap without inventing an untested sixth plane.

Native action risk

The original ApprovalMode values—read_only, local_write, network, delegated, and destructive—are useful operational labels but not a coherent ordering. network describes location, delegated describes authority, and destructive describes reversibility.

The 2026 contract adds ActionRisk:

{
  "effect": "external_state",
  "authority": "user_delegated",
  "reversibility": "compensatable",
  "interaction": "browser",
  "data_scope": "CONFIDENTIAL",
  "decision_ttl_seconds": 120
}

Policy checks each dimension and the actual arguments immediately before execution. The legacy approval mode remains in the v1 ToolCallEnvelope; native risk fields are additive in v1. Making them required in a future major envelope is an evidence-gated migration decision, not a promised release.

Durable and delegated execution

A durable session pins its Context Pack, policy bundles, knowledge snapshot, model-routing policy, evaluator suite, and tool registry view. Resume restores atomic budget usage and continues after the last committed Critic verdict. It never replays a completed effect merely because a process restarted.

A delegation is a derived run, not an informal message. The parent issues a child identity claim with:

  • allowed intents and output schema;
  • capability allow-list and argument ceilings;
  • data classification ceiling;
  • token, cost, tool-call, and wall-clock sub-budgets;
  • expiration and cancellation semantics;
  • a rule that child effects remain isolated until the parent Critic accepts them.

The parent’s authority is an upper bound. Delegation cannot mint permission, reset consumed budget, or hide model/tool calls from the parent trace.

Multimodal and computer-use agents

Screenshots, audio, documents, and UI state enter the Context plane as content-addressed artifacts. Extracted text is derived evidence and retains a pointer to the source artifact and extraction model. Instructions found inside retrieved content or a UI are untrusted data unless policy explicitly promotes them.

Computer-use actions require a pre-action observation, a bounded action proposal, a post-action observation, and an environment receipt. Sensitive inputs are referenced through secret handles, not exposed to the model as plaintext. For irreversible or ambiguous UI actions, the Critic must prefer a human handoff over coordinate guessing.

Adoption path

StageRequired changeCompatibility
0. InventoryList model calls, tools, background jobs, agent peers, browser/computer surfaces, and current approvalsNo runtime change
1. ObserveEmit principal chains, routing lineage, context ledgers, conflict markers, tool-result postcondition/recovery evidence, and replay packetsAdditive telemetry
2. ClassifyAdd ActionRisk to capabilities, validate conservative legacy-mode projections, and map conflicts to evidence requirementsv1 envelopes remain accepted
3. EnforceApply per-dimension policy and just-in-time argument checksMixed v1/native adapters supported
4. DurabilizeAdd checkpoints, artifact pins, cancellation, and monotonic budget restorationSynchronous runs unchanged
5. DelegateIntroduce child claims, sub-budgets, isolated effects, and merge verdictsNo implicit agent-to-agent authority
6. PromoteRequire replay, holdouts, release gates, and rollback for harness changesManual promotion remains valid

What is deliberately not changed

  • The five planes remain the architecture map; new capabilities fit within them.
  • Policy stays outside model reasoning.
  • The Tool Gateway remains the only path to external effects.
  • Durable memory still requires promotion and provenance.
  • DecisionRecord remains the audit receipt, and replay remains the test of operational truth.
  • Protocols such as MCP and A2A describe interoperability; they do not grant authority.

Review checklist

A 2026-ready workflow can answer all of these with typed artifacts:

  • Which principal and delegated authority permitted each effect?
  • What were the action’s effect, reversibility, interaction, and data boundaries?
  • What exact context did each model see, and which candidate context was omitted?
  • Which required claims were conflicted, how were they resolved, and which evidence branch was superseded?
  • Which model profile was selected, why, and which fallbacks were rejected?
  • Can the run resume without refreshing pins or repeating side effects?
  • Did every child agent receive narrower authority and a real sub-budget?
  • Can the decision be replayed with side effects disabled?
  • Can every promoted harness change be rolled back independently?

Success measures

The late-2026 program succeeds when:

  • two or more runtimes pass the same deterministic conformance fixtures;
  • every externally visible effect is linked to policy, identity, idempotency, postcondition, and recovery evidence;
  • native-risk coverage rises without breaking v1 envelope consumers;
  • protocol upgrades are isolated profile changes rather than core-contract rewrites;
  • replay can compare outcomes without re-executing live effects; and
  • proposed core types are rejected as readily as they are added when implementation evidence is weak.

Next steps