Agent engineering series
How strong AI engineers build agents with datasets, scorecards, traces, and harness improvement loops.

Jev and System One Models: What Changes When AI Returns Decisions?
A critical engineering review of Jev and System One models: typed probabilistic decisions, calibration, abstention, latency, schema design, governed execution, and the evidence required before automation.
Harness Engineering on September 26, 2026: From Agent Loops to Personal-Agent Operating Systems
A current practitioner field report on harness engineering through September 26, 2026: Meta Muse, JAZ’s minimalist harness, event-sourced coding agents, agentic-commerce access disputes, embodied agents, and the next research questions.
Harness Engineering: The September 2026 Practitioner’s Field Guide
A primary-source review through September 15, 2026: controlled harness comparisons, evolving safety defenses, context inheritance, managed runtimes, implementation availability, and the next experiments.
Subagent Context Inheritance and Independent Review
A practitioner’s design and evaluation guide for forked and isolated subagents: context selection, reviewer independence, cache economics, information boundaries, and a reproducible experiment plan.
Managed Agent Harnesses: The Control Boundaries You Still Own
A research-backed operating guide to managed agent harnesses: durable sessions, caller credentials, authorization, interrupted effects, version drift, migration tests, and application-owned evidence.
Harness Engineering in September 2026: Measurement, Evolution, and Operational Reality
A critical research map of harness engineering through September 10, 2026: new benchmarks, automatic optimization, reusable tools, long-running agents, public implementations, and unresolved questions.
FrontierHarness Eval: Auditing the 17.5× Cost Gap
A reproducible audit of FrontierHarness Eval: cost per success versus median cost, paired task outcomes, small-sample uncertainty, and what to measure before choosing a production harness.
Self-Improving Harnesses: HarnessDev, HoH, and the Promotion Boundary
A practitioner analysis of HarnessDev, Harness-of-Harness, HarnessOpt-Bench, and HarnessCompass, with concrete boundaries for evaluation, durable state, recovery, and harness promotion.
A Research and Publishing Roadmap for Harness Engineering
A prioritized roadmap for original harness engineering research and technical publishing, with experiments, baselines, artifacts, decision gates, and falsifiable directions through 2027.
Agent Harness Benchmark: An Experimental Protocol Beyond Model Scores
A proposed benchmark protocol for isolating harness effects, injecting operational faults, measuring repeated reliability, and publishing auditable scorecards.
The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes
A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.
Skill Lift: How to Gate Agent Skills With Paired Live Evals
An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.
Harness Engineering in August 2026: The Control Plane Gets Measured
A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.
DeepSeek Harness: Everything Is a Plugin—Except Trust
A source-code critique of DeepSeek Harness and Cordis: why an everything-is-a-plugin runtime matters, where reversibility stops, and why self-evolution still needs an immutable control plane.

Agentic Recommender Systems: How Harness Engineering Changes Personalization
Recommenders are evolving from rankers into tool-using agents. Here is the hybrid architecture, evaluation harness, and rollout plan needed to build them safely.
Prompt Engineering and LLM Latency: The 10-Day Field Report
A source-backed field report on the last ten days of prompt engineering and LLM latency: OpenAI Ultrafast and long-context Fast mode, vLLM and SGLang releases, new inference research, production best practices, and the roadmap ahead.
Production Prompt Engineering in 2026: From Instructions to an Evaluated Contract
A deep production guide to prompt engineering: lean instruction design, stable cacheable prefixes, tool and output contracts, reasoning budgets, datasets, robustness tests, versioning, rollout gates, and a 90-day implementation roadmap.
LLM Latency Engineering: TTFT, Caching, Routing, and the Road to Real-Time Agents
A production guide to LLM latency optimization across queueing, prefill, decode, output budgets, prompt caching, request topology, model and service-tier routing, speculative decoding, disaggregated serving, observability, and rollout.
Harness Engineering: The Runtime Became the Product
A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.
Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation
A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.
The Production Harness Roadmap: Lessons from Claude Code, Copilot Studio, and Thea
A concrete 90-day roadmap for building a governed agent harness, derived from the latest Claude Code, Copilot Studio, Thea, Harness-IF, and A2E evidence.

Agent Evaluation After the Leaderboard: The KDD 2026 Production Assurance Playbook
KDD 2026 shows why agent evaluation must cover protocols, trajectories, external state, evaluator operating profiles, causal rollout, and drift.

Multi-Agent Consensus Is Not Correctness: How Debate Manufactures Confidence
Why agent agreement can hide correlated error—and how to build governed deliberation with effective agent count, independent verification, and release gates.

OpenWorker Review: A Real Desktop Coworker With an Unfinished Trust Runtime
OpenWorker already owns the agent loop, approvals, connectors, and desktop UX. Its next leap is containment, durable effects, replay, budgets, and evals.

Graph Engineering for AI Agents: The X Hype Cycle and a Production Playbook
Separate graph-engineering signal from X hype, then compile agent workflows with governed transitions, graph invariants, durable effects, and exact replay.

Reverse Prompting: Work Backward from the Output, Then Prove the Prompt
Reverse prompting is useful when it turns a desired output into a candidate instruction contract and validates that contract forward. Here is the research, the workflow, the failure modes, and the security boundary.

Passive Awareness Is the Missing Multi-Agent Primitive—But It Needs a Control Plane
AgentRadio shows that agents improve when discoveries can cross lane boundaries during execution. Production teams should keep the asynchronous channel, then add typed evidence, bounded delivery, authority isolation, and replay.

Your Agent Benchmark Is Measuring the Harness
Two new evaluation systems separate benchmark, harness, model, and environment. Their results show why one leaderboard score cannot tell you which agent to ship.

Multi-Agent Systems Need Replay Before More Agents
Multi-agent systems multiply coordination, cost, and causal ambiguity. Prove one agent is replayable before adding lanes and handoffs.

How to Develop an Agent with an Agent Harness, End to End
An end-to-end field guide for building agents as measurable harnesses: context, planning, tools, records, evals, rollout, and learning.

How Great AI Engineers Build Agents: Datasets, Scores, and Harnesses That Improve
Why strong AI engineers build datasets, scorecards, traces, and improvement loops instead of treating agents as prompts plus tools.

Dataset-First Agent Engineering: The Golden Sets Behind Reliable Agents
A practical guide to golden sets, task distributions, corrected runs, held-out releases, and production slices for agent engineering.

Scorecards Over Vibes: The Five Metrics That Keep Agents Honest
The five metrics that keep agents honest: policy, utility, latency, safety, and economics.

Trace Review Is the Agent Debugger: Grade the Path, Not Just the Answer
How trace review grades the path, not just the answer, by inspecting context, plans, tools, guardrails, critic verdicts, and corrections.

Harness Candidates Are Model Checkpoints: How to Improve Agents Without Silent Mutation
How to treat every prompt, retrieval, tool, policy, and evaluator change as a scored, reviewed, reversible harness candidate.