Skip to content

Blog

Operator-grade essays on harness engineering, context packs, replay, evaluators, and running agents in production.

Get new posts by email

ContextOS essays, field notes, and implementation guides.

RSS remains available for feed readers.

Start with the path that matches your job

The best ContextOS essays are built as field guides: pick the role path, then go deeper through the series.

Open start guide

Start here

If you have read nothing yet, read these in order.

OpenAI 2026 research series

A source-backed field guide to OpenAI's latest models, agents, developer platform, multimodal systems, science, safety, and infrastructure.

August 1, 2026·7 min

OpenAI in 2026: A Full-Stack Map of Its Latest Work

A source-backed map of OpenAI's newest work across models, agents, multimodal products, science, safety, and compute—and how the pieces fit.

August 7, 2026·12 min

OpenAI's Latest Work: A 10-Day Field Report

A source-backed 10-day delta across Codex, plugins, API operations, model economics, security scanning, and the GPT-5.4 migration.

August 1, 2026·6 min

GPT-5.6 Field Guide: Models, Tools, Reasoning, and the New Runtime

A practical guide to GPT-5.6 Sol, Terra, and Luna—and the runtime features that matter more than a headline benchmark score.

August 1, 2026·6 min

ChatGPT Work, Codex, and OpenAI Presence: The Emerging Agent Stack

How ChatGPT Work, Codex, and OpenAI Presence form distinct layers of OpenAI's shift from helpful answers to governed, long-running work.

August 1, 2026·6 min

Inside OpenAI's 2026 Agent Platform: Responses, Tools, MCP, and Control

A builder's map of the Responses API, programmatic tools, multi-agent orchestration, MCP, identity, spend controls, and deployment boundaries.

August 1, 2026·6 min

OpenAI's Multimodal Stack: GPT-Live, GPT Image 2, and Sora 2

A product and developer guide to OpenAI's latest voice, image, and video systems—and the boundaries between interaction and media production.

August 1, 2026·6 min

OpenAI's Latest Science Work: From Benchmarks to Verifiable Discovery

What OpenAI's newest mathematics, scientific-computing, biology, and evaluation work says about AI as a research collaborator.

August 1, 2026·7 min

OpenAI's Frontier Stack: Long-Horizon Safety, Stargate, and Custom Silicon

Why OpenAI is scaling trajectory-level safeguards, cyber containment, data centers, and custom inference hardware alongside frontier models.

Anthropic 2026 research series

A source-backed field guide to Anthropic's latest Claude models, agent products, developer platform, research, safety, and company strategy.

August 4, 2026·10 min

Anthropic in 2026: A Full-Stack Map of Its Latest Work

A source-backed map of Anthropic's latest work across Claude models, agents, the developer platform, safety, research, enterprise strategy, and compute.

August 7, 2026·14 min

Anthropic's Latest Work: Safety Moves Into the Runtime

A source-backed field report on Fable 5's biology safeguard update, enterprise inference hooks, skill and plugin scanning, Claude Code security, and Anthropic's global policy expansion.

August 4, 2026·8 min

Claude 5 Field Guide: Fable, Mythos, Opus, Sonnet, and Haiku

A practical guide to Anthropic's current Claude lineup, including capability tiers, pricing, context, effort, safeguards, fallbacks, and model-selection tradeoffs.

August 4, 2026·8 min

From Claude Code to Cowork and Claude Tag: Anthropic's Agent Product Stack

How Claude evolved from chat to coding, cross-app work, and a persistent team agent—and what identity, memory, permissions, and oversight each surface requires.

August 4, 2026·9 min

Inside Anthropic's 2026 Claude Platform: Managed Agents, Tools, Memory, and MCP

A builder's map of the Messages API, Agent SDK, Managed Agents, programmatic tools, MCP, Skills, memory, task budgets, and enterprise controls.

August 4, 2026·10 min

Anthropic's 2026 Safety Stack: Alignment, Safeguards, Containment, and Disclosure

A critical map of Anthropic's latest safety work across constitutional training, capability evaluations, classifiers, model routing, containment, trusted access, and incident disclosure.

August 4, 2026·10 min

Anthropic's Latest Research: Model Minds, Scientific Agents, Robotics, and the Economy

A source-backed review of Anthropic's newest work in interpretability, alignment, cryptanalysis, biology, robotics, autonomous systems, values, and economic measurement.

August 4, 2026·9 min

Anthropic's 2026 Business Strategy: Compute, Enterprise Distribution, and Governance

How Anthropic is financing multi-cloud compute, distributing Claude through enterprises and partners, and using safety governance as both mission and market position.

AI literacy series

Mental models for business leaders, domain experts, and operators learning how to think about real agentic systems.

The AI Agent Accountability Matrix: Who Owns a Failed Decision? illustration
July 20, 2026·15 min

The AI Agent Accountability Matrix: Who Owns a Failed Decision?

AI agent failures cross models, tools, data, approvals, and operations. Use this accountability matrix to assign owners before the incident.

AI Tokenomics: From Cost per Token to Cost per Trusted Outcome illustration
May 26, 2026·16 min

AI Tokenomics: From Cost per Token to Cost per Trusted Outcome

AI tokenomics connects cost per token, agentic cost multipliers, routing, evals, governance, and cost per trusted outcome.

The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do illustration
May 23, 2026·12 min

The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do

A practical governance model for granting AI agents bounded authority based on risk, evidence, policy confidence, evals, and approval.

AI Agents for Business Leaders: Build the Airport, Not Just the Plane illustration
May 13, 2026·20 min

AI Agents for Business Leaders: Build the Airport, Not Just the Plane

A practical executive playbook for agentic AI: define the work, evidence, authority, scorecards, approvals, security, observability, and improvement loop.

Before Your Team Asks for an AI Agent, Map the Real Work illustration
May 13, 2026·4 min

Before Your Team Asks for an AI Agent, Map the Real Work

A practical guide for business teams mapping real work before building agents: actors, evidence, tools, decisions, risks, exceptions, and feedback loops.

Trusting AI at Work: Approvals, Boundaries, and Receipts illustration
May 13, 2026·4 min

Trusting AI at Work: Approvals, Boundaries, and Receipts

A plain-English guide to agent trust: what AI can read, draft, send, change, approve, and how receipts make decisions accountable.

How to Judge AI Work: Scorecards, Not Vibes illustration
May 13, 2026·4 min

How to Judge AI Work: Scorecards, Not Vibes

A practical guide for business teams evaluating AI agents with scorecards, examples, traces, human corrections, and launch gates instead of demos and vibes.

AI Does Not Launch Once: Feedback Loops After Go-Live illustration
May 13, 2026·8 min

AI Does Not Launch Once: Feedback Loops After Go-Live

A plain-English guide to operating agents after launch: corrections, recurring failures, proposal queues, rollout, rollback, and review.

AI agents in the real world

Research-backed field guides to the delegation, verification, permission, memory, and exception failures hidden by capability demos.

Product management series

How product managers shape real agentic systems with intents, authority, scorecards, rollout gates, and improvement loops.

Agent engineering series

How strong AI engineers build agents with datasets, scorecards, traces, and harness improvement loops.

25 postsOpen series
August 24, 2026·13 min

The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes

A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.

August 24, 2026·12 min

Skill Lift: How to Gate Agent Skills With Paired Live Evals

An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.

August 24, 2026·16 min

Harness Engineering in August 2026: The Control Plane Gets Measured

A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.

August 16, 2026·15 min

DeepSeek Harness: Everything Is a Plugin—Except Trust

A source-code critique of DeepSeek Harness and Cordis: why an everything-is-a-plugin runtime matters, where reversibility stops, and why self-evolution still needs an immutable control plane.

Agentic Recommender Systems: How Harness Engineering Changes Personalization illustration
August 15, 2026·26 min

Agentic Recommender Systems: How Harness Engineering Changes Personalization

Recommenders are evolving from rankers into tool-using agents. Here is the hybrid architecture, evaluation harness, and rollout plan needed to build them safely.

August 14, 2026·14 min

Prompt Engineering and LLM Latency: The 10-Day Field Report

A source-backed field report on the last ten days of prompt engineering and LLM latency: OpenAI Ultrafast and long-context Fast mode, vLLM and SGLang releases, new inference research, production best practices, and the roadmap ahead.

August 14, 2026·12 min

Production Prompt Engineering in 2026: From Instructions to an Evaluated Contract

A deep production guide to prompt engineering: lean instruction design, stable cacheable prefixes, tool and output contracts, reasoning budgets, datasets, robustness tests, versioning, rollout gates, and a 90-day implementation roadmap.

August 14, 2026·16 min

LLM Latency Engineering: TTFT, Caching, Routing, and the Road to Real-Time Agents

A production guide to LLM latency optimization across queueing, prefill, decode, output budgets, prompt caching, request topology, model and service-tier routing, speculative decoding, disaggregated serving, observability, and rollout.

August 13, 2026·13 min

Harness Engineering: The Runtime Became the Product

A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.

August 13, 2026·12 min

Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation

A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.

August 13, 2026·13 min

The Production Harness Roadmap: Lessons from Claude Code, Copilot Studio, and Thea

A concrete 90-day roadmap for building a governed agent harness, derived from the latest Claude Code, Copilot Studio, Thea, Harness-IF, and A2E evidence.

Agent Evaluation After the Leaderboard: The KDD 2026 Production Assurance Playbook illustration
August 12, 2026·17 min

Agent Evaluation After the Leaderboard: The KDD 2026 Production Assurance Playbook

KDD 2026 shows why agent evaluation must cover protocols, trajectories, external state, evaluator operating profiles, causal rollout, and drift.

Multi-Agent Consensus Is Not Correctness: How Debate Manufactures Confidence illustration
August 12, 2026·15 min

Multi-Agent Consensus Is Not Correctness: How Debate Manufactures Confidence

Why agent agreement can hide correlated error—and how to build governed deliberation with effective agent count, independent verification, and release gates.

OpenWorker Review: A Real Desktop Coworker With an Unfinished Trust Runtime illustration
August 4, 2026·16 min

OpenWorker Review: A Real Desktop Coworker With an Unfinished Trust Runtime

OpenWorker already owns the agent loop, approvals, connectors, and desktop UX. Its next leap is containment, durable effects, replay, budgets, and evals.

Graph Engineering for AI Agents: The X Hype Cycle and a Production Playbook illustration
July 31, 2026·25 min

Graph Engineering for AI Agents: The X Hype Cycle and a Production Playbook

Separate graph-engineering signal from X hype, then compile agent workflows with governed transitions, graph invariants, durable effects, and exact replay.

Reverse Prompting: Work Backward from the Output, Then Prove the Prompt illustration
July 31, 2026·12 min

Reverse Prompting: Work Backward from the Output, Then Prove the Prompt

Reverse prompting is useful when it turns a desired output into a candidate instruction contract and validates that contract forward. Here is the research, the workflow, the failure modes, and the security boundary.

Passive Awareness Is the Missing Multi-Agent Primitive—But It Needs a Control Plane illustration
July 31, 2026·10 min

Passive Awareness Is the Missing Multi-Agent Primitive—But It Needs a Control Plane

AgentRadio shows that agents improve when discoveries can cross lane boundaries during execution. Production teams should keep the asynchronous channel, then add typed evidence, bounded delivery, authority isolation, and replay.

Your Agent Benchmark Is Measuring the Harness illustration
July 24, 2026·10 min

Your Agent Benchmark Is Measuring the Harness

Two new evaluation systems separate benchmark, harness, model, and environment. Their results show why one leaderboard score cannot tell you which agent to ship.

Multi-Agent Systems Need Replay Before More Agents illustration
July 19, 2026·11 min

Multi-Agent Systems Need Replay Before More Agents

Multi-agent systems multiply coordination, cost, and causal ambiguity. Prove one agent is replayable before adding lanes and handoffs.

How to Develop an Agent with an Agent Harness, End to End illustration
May 12, 2026·17 min

How to Develop an Agent with an Agent Harness, End to End

An end-to-end field guide for building agents as measurable harnesses: context, planning, tools, records, evals, rollout, and learning.

How Great AI Engineers Build Agents: Datasets, Scores, and Harnesses That Improve illustration
May 12, 2026·13 min

How Great AI Engineers Build Agents: Datasets, Scores, and Harnesses That Improve

Why strong AI engineers build datasets, scorecards, traces, and improvement loops instead of treating agents as prompts plus tools.

Dataset-First Agent Engineering: The Golden Sets Behind Reliable Agents illustration
May 12, 2026·7 min

Dataset-First Agent Engineering: The Golden Sets Behind Reliable Agents

A practical guide to golden sets, task distributions, corrected runs, held-out releases, and production slices for agent engineering.

Scorecards Over Vibes: The Five Metrics That Keep Agents Honest illustration
May 12, 2026·6 min

Scorecards Over Vibes: The Five Metrics That Keep Agents Honest

The five metrics that keep agents honest: policy, utility, latency, safety, and economics.

Trace Review Is the Agent Debugger: Grade the Path, Not Just the Answer illustration
May 12, 2026·6 min

Trace Review Is the Agent Debugger: Grade the Path, Not Just the Answer

How trace review grades the path, not just the answer, by inspecting context, plans, tools, guardrails, critic verdicts, and corrections.

Harness Candidates Are Model Checkpoints: How to Improve Agents Without Silent Mutation illustration
May 12, 2026·6 min

Harness Candidates Are Model Checkpoints: How to Improve Agents Without Silent Mutation

How to treat every prompt, retrieval, tool, policy, and evaluator change as a scored, reviewed, reversible harness candidate.

Agent security engineering series

Threat-model, contain, authorize, and red-team tool-using agents with deterministic controls around the model.

10 postsOpen series
August 24, 2026·11 min

HarnessRisk in Practice: Turn the Agent Lifecycle Into a Security Test Matrix

A deep implementation guide to lifecycle security testing for agent harnesses: configuration, extensions, runtime, persistence, actions, recovery, evidence oracles, and release gates.

Persistent Memory Poisoning: The Attack That Outlives the Session illustration
August 12, 2026·16 min

Persistent Memory Poisoning: The Attack That Outlives the Session

A production security architecture for attributable, revalidated, authority-bounded, traceable, and selectively reversible agent memory.

The AI Agent Access Graph: What CISOs Need to See illustration
July 19, 2026·14 min

The AI Agent Access Graph: What CISOs Need to See

AI agents compose identities, tools, data, and delegated authority. Build the access graph, drift alerts, and revocation controls your CISO needs.

Threat-Model an AI Agent: Sources, Sinks, Authority, and Blast Radius illustration
July 12, 2026·9 min

Threat-Model an AI Agent: Sources, Sinks, Authority, and Blast Radius

A practical AI agent threat-modeling method that maps untrusted sources to dangerous sinks, then constrains identity, authority, data, and blast radius at deterministic runtime boundaries.

Prompt Injection Is a Boundary Problem, Not a Prompt Problem illustration
February 21, 2026·9 min

Prompt Injection Is a Boundary Problem, Not a Prompt Problem

Why "smarter prompts" don't defend against indirect prompt injection, and what changes when authority lives outside the model's view.

Agent Identity Is the New Trust Boundary illustration
May 17, 2026·13 min

Agent Identity Is the New Trust Boundary

A practical model for separating agent identity, workload proof, user delegation, scoped authority, and audit across MCP and A2A.

Secure the MCP and Tool Supply Chain: Trust Must Be Continuous illustration
July 12, 2026·15 min

Secure the MCP and Tool Supply Chain: Trust Must Be Continuous

MCP tool descriptions enter the agent's decision loop. Use this attack trace, control scorecard, and policy to stop metadata from becoming authority.

Build the Tool Gateway: The Boundary That Actually Stops a Bad Action illustration
May 5, 2026·5 min

Build the Tool Gateway: The Boundary That Actually Stops a Bad Action

A build-along for the Tool Gateway: adapter manifests, typed envelopes, resolver checks, dispatch, and destructive-action boundaries.

Approval Gates in Code: The Destructive-Mode Handshake illustration
May 6, 2026·5 min

Approval Gates in Code: The Destructive-Mode Handshake

A build-along for approval gates: frozen evidence, human signatures, gateway redemption, and replayable destructive-action handshakes.

Red-Team Agent Hijacking: Build a Security Eval Gate for Repeat Attacks illustration
July 12, 2026·9 min

Red-Team Agent Hijacking: Build a Security Eval Gate for Repeat Attacks

A practical agent-hijacking evaluation harness: scenario design, adaptive and repeated attempts, path-aware metrics, deterministic release gates, and production replay.

Architecture & foundations

The five planes, why prompts alone do not scale, what context engineering means.

11 postsOpen series
Proof-Carrying Context: Why AI Agents Need More Than a Context Window illustration
August 28, 2026·17 min

Proof-Carrying Context: Why AI Agents Need More Than a Context Window

A research-grounded revision of context engineering: from relevant tokens to a replayable view of evidence, conflicts, omissions, and decision sufficiency.

The Glass Runtime: Keeping Humans Close to the Material in an Agentic World illustration
August 11, 2026·15 min

The Glass Runtime: Keeping Humans Close to the Material in an Agentic World

As AI makes output abundant, the central design challenge shifts to preserving human judgment, understanding, intervention, and agency. The Glass Runtime is an architecture for progressive autonomy, inspectable decisions, material-native interaction, reversibility, governed authority, and accountable learning.

The State of AI Agents in 2026: Standards Converged, Models Improved, Production Moved to the Harness illustration
June 21, 2026·13 min

The State of AI Agents in 2026: Standards Converged, Models Improved, Production Moved to the Harness

A mid-2026 review of agentic AI: MCP, A2A and AP2 converged as standards and models got more reliable — yet the bottleneck moved to the governed agent harness.

SecondBrain: A Local-First Agent Operating System You Can Run, Inspect, and Trust illustration
June 26, 2026·8 min

SecondBrain: A Local-First Agent Operating System You Can Run, Inspect, and Trust

SecondBrain is an open-source, local-first agent OS: cognition, memory, governed tools, durable sessions, workflows, and bounded self-improvement in one inspectable runtime you run on your own machine. Here is how it works and how to run it.

Agent Harness: An Architectural Framework for Production AI Agents illustration
May 19, 2026·33 min

Agent Harness: An Architectural Framework for Production AI Agents

A whitepaper on typed contracts, policy gates, traces, verification loops, and release control for production AI agents.

Antahkarana Stack: A Cognitive Layer for Local-First Agents illustration
May 20, 2026·17 min

Antahkarana Stack: A Cognitive Layer for Local-First Agents

A builder-facing explanation of Antahkarana as an engineering layer inspired by the inner faculties of Manas, Buddhi, Chitta, and Ahamkara.

The Five Planes of Agentic Operating Systems illustration
April 29, 2026·10 min

The Five Planes of Agentic Operating Systems

A working decomposition for production agent systems: Intelligence, Context, Decision, Action, and Trust.

ContextOS: A Research-Grounded Architecture for Governed Agent Runtimes illustration
May 16, 2026·28 min

ContextOS: A Research-Grounded Architecture for Governed Agent Runtimes

A research-grounded framing of ContextOS as a governed runtime for context, tools, memory, security, evaluation, replay, and optimization.

Beyond Prompts: The Architecture of Trust for Agentic AI illustration
March 2, 2026·14 min

Beyond Prompts: The Architecture of Trust for Agentic AI

Building a governed decision runtime across Intelligence, Context, Decision, Action, and Trust — with evaluator scoring, approval tiers, and replay-bound audit.

Context Engineering in Production illustration
February 4, 2026·8 min

Context Engineering in Production

Why most agent failures are not model failures — they are context failures — and what changes when context becomes a versioned, testable, replayable contract.

Context Packs in Practice: From Spec to Run illustration
March 26, 2026·8 min

Context Packs in Practice: From Spec to Run

A practical walkthrough of Context Packs: buckets, policy bundles, evaluation gates, lifecycle, and the compile pipeline.

Building the runtime

Compile, gateway, Critic, evaluators, failure handling — the per-request pipeline.

Trust, audit, governance

Replay, approval modes, approval-gate handshakes, and the security boundary.

Memory & evidence

How agents remember, what gets promoted, how knowledge is grounded.

Agent Memory Should Be Compiled at Recall, Not Replayed From Storage illustration
July 31, 2026·13 min

Agent Memory Should Be Compiled at Recall, Not Replayed From Storage

What MemHarness teaches us about reconstruction, negative transfer, source state, and compiling production agent memory against current evidence, policy, and authority.

Agent Memory Is a Systems Workload: What SelfMem Changes—and What It Does Not illustration
July 24, 2026·9 min

Agent Memory Is a Systems Workload: What SelfMem Changes—and What It Does Not

SelfMem improves long-horizon recall by letting an agent optimize its memory strategy. New systems and security research shows the production contract must also cover cost, freshness, provenance, and poisoning.

AI Agent Memory Is Broken: Designing Multi-Layer Memory for Production AI Agents illustration
June 2, 2026·31 min

AI Agent Memory Is Broken: Designing Multi-Layer Memory for Production AI Agents

A production guide to AI agent memory architecture: designing long-term memory for AI agents across working, episodic, semantic, procedural, and organizational layers. Why RAG is not memory, why vector databases are not memory, and how governed, situation-aware memory prevents memory poisoning in enterprise AI agents.

Give Claude Code, Cursor, and Codex Persistent, Auditable Memory illustration
June 27, 2026·6 min

Give Claude Code, Cursor, and Codex Persistent, Auditable Memory

Coding agents are brilliant and amnesiac. SecondBrain's open-source Memory API gives Claude Code, Cursor, Codex, and ChatGPT shared, local-first memory over HTTP and MCP — where every answer carries a citation back to the source chunk.

The Identity Layer: Agents Need Two Identities, Not One illustration
May 14, 2026·6 min

The Identity Layer: Agents Need Two Identities, Not One

Why governed agent runs need entity identity, delegated user identity, and workload identity in the same RunContext.

Promotion-Aware Memory: Capture, Review, Promote, Recall in Code illustration
April 25, 2026·6 min

Promotion-Aware Memory: Capture, Review, Promote, Recall in Code

A build-along for agent memory: capture, review, promote, recall, contradiction checks, and governed memory writes.

Context Graphs: Decision Lineage as a System of Record illustration
April 18, 2026·8 min

Context Graphs: Decision Lineage as a System of Record

How hash-chained DecisionRecords turn execution-time context into a queryable lineage graph for why an agent acted.

Enterprise use cases

Concrete agentic workflows for incident response, financial crime, regulated operations, data stewardship, and software delivery.

Reviewers & improvement

Reviewer agents, rollouts, operator corrections becoming versioned StrategyRules.

What Production Agent Runtimes Actually Teach ContextOS: Twelve Laws of a Governed Harness illustration
July 25, 2026·24 min

What Production Agent Runtimes Actually Teach ContextOS: Twelve Laws of a Governed Harness

The strongest production agent runtimes converge on twelve architectural laws: compile context, persist state, separate authority from containment, make side effects resumable, treat approvals as typed interrupts, and promote learning only through evidence and replay.

Adaptive Agent Harnesses: Learn From Experience Without Letting Production Rewrite Itself illustration
July 24, 2026·10 min

Adaptive Agent Harnesses: Learn From Experience Without Letting Production Rewrite Itself

MemoHarness shows that execution experience can improve the control layer around an LLM. Production systems need a stricter pattern: adaptive performance inside an immutable safety envelope.

Harness Improvement Loops Need Replayable Environments illustration
May 14, 2026·7 min

Harness Improvement Loops Need Replayable Environments

Why harness improvement needs replayable episodes, bounded mutations, scorecards, source closure, and promotion gates.

Autotune the Harness: Baking the Improvement Loop into ContextOS illustration
May 12, 2026·11 min

Autotune the Harness: Baking the Improvement Loop into ContextOS

How ContextOS treats autotune as a gated loop over traces, scorecards, replay sets, bounded candidates, approval, and rollout.

Building a Compliance Reviewer Agent in 60 Lines and a Golden Set illustration
March 15, 2026·6 min

Building a Compliance Reviewer Agent in 60 Lines and a Golden Set

How to build a compliance reviewer agent with a typed verdict envelope, rubric, golden set, and change-control queue.

Building a Reliability Reviewer Agent: 70 Lines Past the Compliance One illustration
March 18, 2026·4 min

Building a Reliability Reviewer Agent: 70 Lines Past the Compliance One

How to extend the reviewer pattern for reliability: timeouts, retries, idempotency, fallback behavior, and rollback declarations.

Pack Rollout in Five Stages: Shipping a Context Pack Without Blowing Up Production illustration
April 11, 2026·7 min

Pack Rollout in Five Stages: Shipping a Context Pack Without Blowing Up Production

A five-stage rollout model for Context Packs: shadow, internal, low-risk, monitored expansion, full release, and rollback.

From Operator Correction to Released StrategyRule: The Improvement Loop, Coded illustration
April 15, 2026·7 min

From Operator Correction to Released StrategyRule: The Improvement Loop, Coded

How one operator correction becomes a reviewed, replayed, versioned StrategyRule that prevents repeat agent failures.