Agent Evaluation
14 essays tagged with Agent Evaluation.
After Meta Muse: A Research and Publishing Roadmap for Personal-Agent Harnesses
A falsifiable research roadmap for personal-agent harnesses: authorization binding, ambiguous-effect recovery, credential misuse, governed memory, counterparty permission, interfaces, deployment, and reasoning architecture.
Harness Engineering on September 26, 2026: From Agent Loops to Personal-Agent Operating Systems
A current practitioner field report on harness engineering through September 26, 2026: Meta Muse, JAZ’s minimalist harness, event-sourced coding agents, agentic-commerce access disputes, embodied agents, and the next research questions.
Jev and System One Models: What Changes When AI Returns Decisions?
A critical engineering review of Jev and System One models: typed probabilistic decisions, calibration, abstention, latency, schema design, governed execution, and the evidence required before automation.
Harness Engineering: The September 2026 Practitioner’s Field Guide
A primary-source review through September 15, 2026: controlled harness comparisons, evolving safety defenses, context inheritance, managed runtimes, implementation availability, and the next experiments.
Subagent Context Inheritance and Independent Review
A practitioner’s design and evaluation guide for forked and isolated subagents: context selection, reviewer independence, cache economics, information boundaries, and a reproducible experiment plan.
FrontierHarness Eval: Auditing the 17.5× Cost Gap
A reproducible audit of FrontierHarness Eval: cost per success versus median cost, paired task outcomes, small-sample uncertainty, and what to measure before choosing a production harness.
A Research and Publishing Roadmap for Harness Engineering
A prioritized roadmap for original harness engineering research and technical publishing, with experiments, baselines, artifacts, decision gates, and falsifiable directions through 2027.
Harness Engineering in September 2026: Measurement, Evolution, and Operational Reality
A critical research map of harness engineering through September 10, 2026: new benchmarks, automatic optimization, reusable tools, long-running agents, public implementations, and unresolved questions.
Self-Improving Harnesses: HarnessDev, HoH, and the Promotion Boundary
A practitioner analysis of HarnessDev, Harness-of-Harness, HarnessOpt-Bench, and HarnessCompass, with concrete boundaries for evaluation, durable state, recovery, and harness promotion.
The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes
A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.
Skill Lift: How to Gate Agent Skills With Paired Live Evals
An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.
Harness Engineering in August 2026: The Control Plane Gets Measured
A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.
Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation
A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.
Harness Engineering: The Runtime Became the Product
A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.