Context
Only current, verified, in-scope evidence enters the run.
ContextOS is the harness around the model: it compiles what an agent may know, constrains what it may do, verifies what happened, and preserves the proof needed to replay or recover the run.
The same run from the hero stays in view. Every stage below exposes the artifact the runtime produced—not a promise about what an agent usually does.
Only current, verified, in-scope evidence enters the run.
Policy remains outside the prompt and evaluates the live action envelope.
The model proposes. The Tool Gateway validates identity, scope, arguments, and effect risk.
The path is scored against evidence and gates, not against the agent’s confidence.
The gateway verifies external state and preserves a concrete recovery handle.
The run closes only when the six operational questions have typed answers.
The planes are not five products in a catalog. They are one governed run viewed at five boundaries, joined by the same principal, budget, policy pins, and trace.
identity · evidence · memory
admission · compilation · budget
plan · verify · score
tools · effects · recovery
policy · evaluation · replay
Better models expand what agents can attempt. Production systems still need to decide what counts as admissible, complete, safe, and worth shipping.
A plausible answer can hide a bad path. Evaluate instructions, tool choices, evidence, effects, and recovery—not only the final text.
Pin the model, harness, Context Pack, skills, tools, policy, evaluators, and rollback target as one compatible build.
Every handoff carries identity, scope, budget, and approval mode. Authority narrows as work moves across agents and tools.
Postconditions, idempotency, compensation, and reversal tokens make safety survive contact with production systems.