Skip to content
Back to Blog
Tag

Agent Evaluation

14 essays tagged with Agent Evaluation.

September 26, 2026·18 min read·Expert

After Meta Muse: A Research and Publishing Roadmap for Personal-Agent Harnesses

A falsifiable research roadmap for personal-agent harnesses: authorization binding, ambiguous-effect recovery, credential misuse, governed memory, counterparty permission, interfaces, deployment, and reasoning architecture.

Read essay
September 26, 2026·12 min read·Expert

Harness Engineering on September 26, 2026: From Agent Loops to Personal-Agent Operating Systems

A current practitioner field report on harness engineering through September 26, 2026: Meta Muse, JAZ’s minimalist harness, event-sourced coding agents, agentic-commerce access disputes, embodied agents, and the next research questions.

Read essay
September 23, 2026·18 min read·Expert

Jev and System One Models: What Changes When AI Returns Decisions?

A critical engineering review of Jev and System One models: typed probabilistic decisions, calibration, abstention, latency, schema design, governed execution, and the evidence required before automation.

Read essay
September 15, 2026·9 min read·Expert

Harness Engineering: The September 2026 Practitioner’s Field Guide

A primary-source review through September 15, 2026: controlled harness comparisons, evolving safety defenses, context inheritance, managed runtimes, implementation availability, and the next experiments.

Read essay
September 15, 2026·8 min read·Expert

Subagent Context Inheritance and Independent Review

A practitioner’s design and evaluation guide for forked and isolated subagents: context selection, reviewer independence, cache economics, information boundaries, and a reproducible experiment plan.

Read essay
September 10, 2026·8 min read·Expert

FrontierHarness Eval: Auditing the 17.5× Cost Gap

A reproducible audit of FrontierHarness Eval: cost per success versus median cost, paired task outcomes, small-sample uncertainty, and what to measure before choosing a production harness.

Read essay
September 10, 2026·12 min read·Expert

A Research and Publishing Roadmap for Harness Engineering

A prioritized roadmap for original harness engineering research and technical publishing, with experiments, baselines, artifacts, decision gates, and falsifiable directions through 2027.

Read essay
September 10, 2026·9 min read·Expert

Harness Engineering in September 2026: Measurement, Evolution, and Operational Reality

A critical research map of harness engineering through September 10, 2026: new benchmarks, automatic optimization, reusable tools, long-running agents, public implementations, and unresolved questions.

Read essay
September 10, 2026·9 min read·Expert

Self-Improving Harnesses: HarnessDev, HoH, and the Promotion Boundary

A practitioner analysis of HarnessDev, Harness-of-Harness, HarnessOpt-Bench, and HarnessCompass, with concrete boundaries for evaluation, durable state, recovery, and harness promotion.

Read essay
August 24, 2026·13 min read·Expert

The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes

A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.

Read essay
August 24, 2026·12 min read·Expert

Skill Lift: How to Gate Agent Skills With Paired Live Evals

An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.

Read essay
August 24, 2026·16 min read·Expert

Harness Engineering in August 2026: The Control Plane Gets Measured

A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.

Read essay
August 13, 2026·12 min read·Expert

Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation

A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.

Read essay
August 13, 2026·13 min read·Expert

Harness Engineering: The Runtime Became the Product

A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.

Read essay