AI Agents
9 essays tagged with AI Agents.
The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes
A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.
Skill Lift: How to Gate Agent Skills With Paired Live Evals
An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.
HarnessRisk in Practice: Turn the Agent Lifecycle Into a Security Test Matrix
A deep implementation guide to lifecycle security testing for agent harnesses: configuration, extensions, runtime, persistence, actions, recovery, evidence oracles, and release gates.
Harness Engineering in August 2026: The Control Plane Gets Measured
A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.
DeepSeek Harness: Everything Is a Plugin—Except Trust
A source-code critique of DeepSeek Harness and Cordis: why an everything-is-a-plugin runtime matters, where reversibility stops, and why self-evolution still needs an immutable control plane.
Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation
A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.
Harness Engineering: The Runtime Became the Product
A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.
From Claude Code to Cowork and Claude Tag: Anthropic's Agent Product Stack
How Claude evolved from chat to coding, cross-app work, and a persistent team agent—and what identity, memory, permissions, and oversight each surface requires.
The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do
A practical governance model for granting AI agents bounded authority based on risk, evidence, policy confidence, evals, and approval.