Skip to content
Back to Blog
Tag

AI Agents

9 essays tagged with AI Agents.

August 24, 2026·13 min read·Expert

The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes

A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.

Read essay
August 24, 2026·12 min read·Expert

Skill Lift: How to Gate Agent Skills With Paired Live Evals

An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.

Read essay
August 24, 2026·11 min read·Expert

HarnessRisk in Practice: Turn the Agent Lifecycle Into a Security Test Matrix

A deep implementation guide to lifecycle security testing for agent harnesses: configuration, extensions, runtime, persistence, actions, recovery, evidence oracles, and release gates.

Read essay
August 24, 2026·16 min read·Expert

Harness Engineering in August 2026: The Control Plane Gets Measured

A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.

Read essay
August 16, 2026·15 min read·Expert

DeepSeek Harness: Everything Is a Plugin—Except Trust

A source-code critique of DeepSeek Harness and Cordis: why an everything-is-a-plugin runtime matters, where reversibility stops, and why self-evolution still needs an immutable control plane.

Read essay
August 13, 2026·12 min read·Expert

Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation

A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.

Read essay
August 13, 2026·13 min read·Expert

Harness Engineering: The Runtime Became the Product

A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.

Read essay
August 4, 2026·8 min read·Intermediate

From Claude Code to Cowork and Claude Tag: Anthropic's Agent Product Stack

How Claude evolved from chat to coding, cross-app work, and a persistent team agent—and what identity, memory, permissions, and oversight each surface requires.

Read essay
May 23, 2026·12 min read·Intermediate

The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do

A practical governance model for granting AI agents bounded authority based on risk, evidence, policy confidence, evals, and approval.

Read essay