About Piyush Kumar
I am Piyush Kumar, an AI platform builder and systems thinker focused on making AI agents reliable for real-world enterprise work.
Through ContextOS AI, I write about the architecture of governed intelligence systems: agents with memory, context, tools, policies, evaluations, observability, and human oversight.
Helping technology leaders, product leaders, architects, and builders move from impressive AI demos to production-grade agentic systems.
Building governed intelligence systems for the agentic era
My focus is not AI as a chatbot layer. It is AI as a new execution layer for business - one that needs context, memory, tools, policies, evaluations, observability, and governance.
ContextOS AI is my attempt to explain this shift in simple, practical language. The goal is to make complex AI architecture understandable, actionable, and useful for people building serious systems.
What I care about
The next phase of AI will not be won by simply calling larger models. It will be won by organizations that design systems where AI can:
My work
AI Platforms
Reusable foundations for agents, copilots, personalization systems, and decisioning engines.
Agentic AI Architecture
Planners, executors, memory, tools, guardrails, evaluations, and human approval loops.
Personalization and Context
Systems that understand users, journeys, intent, preferences, and situational context.
Data and Decision Infrastructure
Data platforms, knowledge graphs, feature systems, embeddings, telemetry, and business rules.
Evaluation and Observability
Scorecards, traces, simulations, and feedback systems for useful, safe, improving AI.
Why ContextOS AI exists
Most AI discussions focus on models. In production, the model is only one part of the system.
Real enterprise AI needs an operating environment around the model: context assembly, memory retrieval, tool orchestration, policy enforcement, approval workflows, audit trails, cost controls, quality evaluations, and continuous learning loops.
ContextOS AI exists to document that architecture shift.
My perspective
I believe AI agents should not be treated as magic workers. They should be treated as governed digital operators.
Every agent needs a clear contract:
Professional background
Over the last several years, I have worked on large-scale consumer technology platforms across AI, personalization, data infrastructure, marketing technology, customer experience, and digital commerce.
My experience includes building and scaling systems with high traffic, high reliability, and high business impact across personalization, growth, conversational AI, customer experience automation, data platforms, evaluation systems, and AI-native product experiences.
That practical exposure shapes the way I write: less abstract AI hype, more systems that can survive production reality.
Current areas of exploration
ContextOS
A governed runtime where memory, tools, policies, evaluations, and observability form a common intelligence layer.
Agent Harness Engineering
The discipline required to make agents repeatable, measurable, safe, and production-ready.
AI Memory Systems
How long-term memory, short-term context, user preferences, knowledge graphs, and retrieval should work together.
Tokenomics
How enterprises should measure the cost, value, and reliability of AI beyond simple token cost.
Evaluation-first AI
Scorecards, simulation, regression testing, and feedback loops before systems are trusted.
Business Leadership in AI
AI adoption as operating infrastructure, not a collection of isolated tools.
A simple belief
AI will not replace systems thinking. It will reward it.
The organizations that win with AI will be the ones that combine models, data, tools, workflows, policies, and people into coherent systems. That is the future I am interested in building and explaining.
Writing
Essays, frameworks, architecture notes, and implementation-oriented thinking by Piyush Kumar.
Proof-Carrying Context: Why AI Agents Need More Than a Context Window
A research-grounded revision of context engineering: from relevant tokens to a replayable view of evidence, conflicts, omissions, and decision sufficiency.
The Harness Engineering Roadmap for 2026–27: From Agent Loops to Proof-Carrying Runtimes
A source-grounded 18-month roadmap for harness engineering across release manifests, portable trajectories, lifecycle security, skill evaluation, delegated authority, recovery, economics, and governed adaptation.
Skill Lift: How to Gate Agent Skills With Paired Live Evals
An implementation guide to evaluating agent skills as executable release artifacts using paired baselines, Skill Lift, ATIF trajectories, security checks, domain graders, and CI gates.
HarnessRisk in Practice: Turn the Agent Lifecycle Into a Security Test Matrix
A deep implementation guide to lifecycle security testing for agent harnesses: configuration, extensions, runtime, persistence, actions, recovery, evidence oracles, and release gates.
Harness Engineering in August 2026: The Control Plane Gets Measured
A source-backed state of the field on harness engineering: what shipped, what the newest research measured, what practitioners are debating, and the roadmap from agent loops to proof-carrying runtimes.
DeepSeek Harness: Everything Is a Plugin—Except Trust
A source-code critique of DeepSeek Harness and Cordis: why an everything-is-a-plugin runtime matters, where reversibility stops, and why self-evolution still needs an immutable control plane.
Agentic Recommender Systems: How Harness Engineering Changes Personalization
Recommenders are evolving from rankers into tool-using agents. Here is the hybrid architecture, evaluation harness, and rollout plan needed to build them safely.
Prompt Engineering and LLM Latency: The 10-Day Field Report
A source-backed field report on the last ten days of prompt engineering and LLM latency: OpenAI Ultrafast and long-context Fast mode, vLLM and SGLang releases, new inference research, production best practices, and the roadmap ahead.
Production Prompt Engineering in 2026: From Instructions to an Evaluated Contract
A deep production guide to prompt engineering: lean instruction design, stable cacheable prefixes, tool and output contracts, reasoning budgets, datasets, robustness tests, versioning, rollout gates, and a 90-day implementation roadmap.
LLM Latency Engineering: TTFT, Caching, Routing, and the Road to Real-Time Agents
A production guide to LLM latency optimization across queueing, prefill, decode, output budgets, prompt caching, request topology, model and service-tier routing, speculative decoding, disaggregated serving, observability, and rollout.
The Production Harness Roadmap: Lessons from Claude Code, Copilot Studio, and Thea
A concrete 90-day roadmap for building a governed agent harness, derived from the latest Claude Code, Copilot Studio, Thea, Harness-IF, and A2E evidence.
Your Agent Followed the Rule—or Did It? Harness-IF and A2E Change Agent Evaluation
A deep implementation guide to rule-level, surface-aware, trajectory-based agent evaluation using the new Harness-IF and A2E research.
Harness Engineering: The Runtime Became the Product
A source-backed 10-day field report on Harness-IF, A2E, Thea, Microsoft Copilot Studio, and Claude Code—and what their combined evidence changes for agent engineering.
Persistent Memory Poisoning: The Attack That Outlives the Session
A production security architecture for attributable, revalidated, authority-bounded, traceable, and selectively reversible agent memory.
Multi-Agent Consensus Is Not Correctness: How Debate Manufactures Confidence
Why agent agreement can hide correlated error—and how to build governed deliberation with effective agent count, independent verification, and release gates.
Agent Evaluation After the Leaderboard: The KDD 2026 Production Assurance Playbook
KDD 2026 shows why agent evaluation must cover protocols, trajectories, external state, evaluator operating profiles, causal rollout, and drift.
The Glass Runtime: Keeping Humans Close to the Material in an Agentic World
As AI makes output abundant, the central design challenge shifts to preserving human judgment, understanding, intervention, and agency. The Glass Runtime is an architecture for progressive autonomy, inspectable decisions, material-native interaction, reversibility, governed authority, and accountable learning.
Anthropic's Latest Work: Safety Moves Into the Runtime
A source-backed field report on Fable 5's biology safeguard update, enterprise inference hooks, skill and plugin scanning, Claude Code security, and Anthropic's global policy expansion.
OpenAI's Latest Work: A 10-Day Field Report
A source-backed 10-day delta across Codex, plugins, API operations, model economics, security scanning, and the GPT-5.4 migration.
OpenWorker Review: A Real Desktop Coworker With an Unfinished Trust Runtime
OpenWorker already owns the agent loop, approvals, connectors, and desktop UX. Its next leap is containment, durable effects, replay, budgets, and evals.
Anthropic in 2026: A Full-Stack Map of Its Latest Work
A source-backed map of Anthropic's latest work across Claude models, agents, the developer platform, safety, research, enterprise strategy, and compute.
Anthropic's 2026 Business Strategy: Compute, Enterprise Distribution, and Governance
How Anthropic is financing multi-cloud compute, distributing Claude through enterprises and partners, and using safety governance as both mission and market position.
Inside Anthropic's 2026 Claude Platform: Managed Agents, Tools, Memory, and MCP
A builder's map of the Messages API, Agent SDK, Managed Agents, programmatic tools, MCP, Skills, memory, task budgets, and enterprise controls.
Anthropic's Latest Research: Model Minds, Scientific Agents, Robotics, and the Economy
A source-backed review of Anthropic's newest work in interpretability, alignment, cryptanalysis, biology, robotics, autonomous systems, values, and economic measurement.
Anthropic's 2026 Safety Stack: Alignment, Safeguards, Containment, and Disclosure
A critical map of Anthropic's latest safety work across constitutional training, capability evaluations, classifiers, model routing, containment, trusted access, and incident disclosure.
Claude 5 Field Guide: Fable, Mythos, Opus, Sonnet, and Haiku
A practical guide to Anthropic's current Claude lineup, including capability tiers, pricing, context, effort, safeguards, fallbacks, and model-selection tradeoffs.
From Claude Code to Cowork and Claude Tag: Anthropic's Agent Product Stack
How Claude evolved from chat to coding, cross-app work, and a persistent team agent—and what identity, memory, permissions, and oversight each surface requires.
ChatGPT Work, Codex, and OpenAI Presence: The Emerging Agent Stack
How ChatGPT Work, Codex, and OpenAI Presence form distinct layers of OpenAI's shift from helpful answers to governed, long-running work.
GPT-5.6 Field Guide: Models, Tools, Reasoning, and the New Runtime
A practical guide to GPT-5.6 Sol, Terra, and Luna—and the runtime features that matter more than a headline benchmark score.
OpenAI in 2026: A Full-Stack Map of Its Latest Work
A source-backed map of OpenAI's newest work across models, agents, multimodal products, science, safety, and compute—and how the pieces fit.
OpenAI's Multimodal Stack: GPT-Live, GPT Image 2, and Sora 2
A product and developer guide to OpenAI's latest voice, image, and video systems—and the boundaries between interaction and media production.
Inside OpenAI's 2026 Agent Platform: Responses, Tools, MCP, and Control
A builder's map of the Responses API, programmatic tools, multi-agent orchestration, MCP, identity, spend controls, and deployment boundaries.
OpenAI's Frontier Stack: Long-Horizon Safety, Stargate, and Custom Silicon
Why OpenAI is scaling trajectory-level safeguards, cyber containment, data centers, and custom inference hardware alongside frontier models.
OpenAI's Latest Science Work: From Benchmarks to Verifiable Discovery
What OpenAI's newest mathematics, scientific-computing, biology, and evaluation work says about AI as a research collaborator.
Graph Engineering for AI Agents: The X Hype Cycle and a Production Playbook
Separate graph-engineering signal from X hype, then compile agent workflows with governed transitions, graph invariants, durable effects, and exact replay.
Reverse Prompting: Work Backward from the Output, Then Prove the Prompt
Reverse prompting is useful when it turns a desired output into a candidate instruction contract and validates that contract forward. Here is the research, the workflow, the failure modes, and the security boundary.
Agent Memory Should Be Compiled at Recall, Not Replayed From Storage
What MemHarness teaches us about reconstruction, negative transfer, source state, and compiling production agent memory against current evidence, policy, and authority.
Human Oversight for Agent Fleets: Confidence Is Not an Audit Policy
A new audit-allocation paper shows that self-reported confidence can make limited human review worse than random and that tiny audit budgets can become rubber-stamping. Production oversight needs risk gates, stratified random coverage, correlation-aware learning, and a measured non-vacuity test.
Passive Awareness Is the Missing Multi-Agent Primitive—But It Needs a Control Plane
AgentRadio shows that agents improve when discoveries can cross lane boundaries during execution. Production teams should keep the asynchronous channel, then add typed evidence, bounded delivery, authority isolation, and replay.
What Production Agent Runtimes Actually Teach ContextOS: Twelve Laws of a Governed Harness
The strongest production agent runtimes converge on twelve architectural laws: compile context, persist state, separate authority from containment, make side effects resumable, treat approvals as typed interrupts, and promote learning only through evidence and replay.
Adaptive Agent Harnesses: Learn From Experience Without Letting Production Rewrite Itself
MemoHarness shows that execution experience can improve the control layer around an LLM. Production systems need a stricter pattern: adaptive performance inside an immutable safety envelope.
Agent Memory Is a Systems Workload: What SelfMem Changes—and What It Does Not
SelfMem improves long-horizon recall by letting an agent optimize its memory strategy. New systems and security research shows the production contract must also cover cost, freshness, provenance, and poisoning.
Your Agent Benchmark Is Measuring the Harness
Two new evaluation systems separate benchmark, harness, model, and environment. Their results show why one leaderboard score cannot tell you which agent to ship.
The AI Agent Accountability Matrix: Who Owns a Failed Decision?
AI agent failures cross models, tools, data, approvals, and operations. Use this accountability matrix to assign owners before the incident.
The AI Agent Access Graph: What CISOs Need to See
AI agents compose identities, tools, data, and delegated authority. Build the access graph, drift alerts, and revocation controls your CISO needs.
Multi-Agent Systems Need Replay Before More Agents
Multi-agent systems multiply coordination, cost, and causal ambiguity. Prove one agent is replayable before adding lanes and handoffs.
Red-Team Agent Hijacking: Build a Security Eval Gate for Repeat Attacks
A practical agent-hijacking evaluation harness: scenario design, adaptive and repeated attempts, path-aware metrics, deterministic release gates, and production replay.
Threat-Model an AI Agent: Sources, Sinks, Authority, and Blast Radius
A practical AI agent threat-modeling method that maps untrusted sources to dangerous sinks, then constrains identity, authority, data, and blast radius at deterministic runtime boundaries.
Secure the MCP and Tool Supply Chain: Trust Must Be Continuous
MCP tool descriptions enter the agent's decision loop. Use this attack trace, control scorecard, and policy to stop metadata from becoming authority.
The AI Software Delivery Squad: From Ticket to Proof-Carrying Pull Request
A production blueprint for coding agents that scope, patch, test, review, and open pull requests without inheriting merge or deploy authority.
Give Claude Code, Cursor, and Codex Persistent, Auditable Memory
Coding agents are brilliant and amnesiac. SecondBrain's open-source Memory API gives Claude Code, Cursor, Codex, and ChatGPT shared, local-first memory over HTTP and MCP — where every answer carries a citation back to the source chunk.
SecondBrain: A Local-First Agent Operating System You Can Run, Inspect, and Trust
SecondBrain is an open-source, local-first agent OS: cognition, memory, governed tools, durable sessions, workflows, and bounded self-improvement in one inspectable runtime you run on your own machine. Here is how it works and how to run it.
The State of AI Agents in 2026: Standards Converged, Models Improved, Production Moved to the Harness
A mid-2026 review of agentic AI: MCP, A2A and AP2 converged as standards and models got more reliable — yet the bottleneck moved to the governed agent harness.
AI Agent Memory Is Broken: Designing Multi-Layer Memory for Production AI Agents
A production guide to AI agent memory architecture: designing long-term memory for AI agents across working, episodic, semantic, procedural, and organizational layers. Why RAG is not memory, why vector databases are not memory, and how governed, situation-aware memory prevents memory poisoning in enterprise AI agents.
Reversibility Is the Missing Safety Primitive for AI Agents
Prevention decides whether agents may act. Reversibility lets them survive being wrong through reversal contracts, compensation, and blast-radius caps.
AI Tokenomics: From Cost per Token to Cost per Trusted Outcome
AI tokenomics connects cost per token, agentic cost multipliers, routing, evals, governance, and cost per trusted outcome.
The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do
A practical governance model for granting AI agents bounded authority based on risk, evidence, policy confidence, evals, and approval.
Antahkarana Stack: A Cognitive Layer for Local-First Agents
A builder-facing explanation of Antahkarana as an engineering layer inspired by the inner faculties of Manas, Buddhi, Chitta, and Ahamkara.
Agent Harness: An Architectural Framework for Production AI Agents
A whitepaper on typed contracts, policy gates, traces, verification loops, and release control for production AI agents.
Agent Identity Is the New Trust Boundary
A practical model for separating agent identity, workload proof, user delegation, scoped authority, and audit across MCP and A2A.
ContextOS: A Research-Grounded Architecture for Governed Agent Runtimes
A research-grounded framing of ContextOS as a governed runtime for context, tools, memory, security, evaluation, replay, and optimization.
Agentic Incident Command Center: Agents Can Coordinate, Boundaries Still Decide
How incident-response agents can coordinate signal, diagnosis, remediation, communications, and approvals without bypassing operational boundaries.
AI Gateway and LLM Router: Model Choice Is a Runtime Decision
How an AI Gateway and LLM Router make model choice policy-bound, budgeted, observable, and replayable across production agent workflows.
Financial Crime Operations: Agentic AI Needs Evidence, Not Autonomy
How KYC, AML, sanctions, and fraud casework can use agentic workflows while preserving evidence, policy gates, and human adjudication.
The Identity Layer: Agents Need Two Identities, Not One
Why governed agent runs need entity identity, delegated user identity, and workload identity in the same RunContext.
MCP Adapters in Production: The Manifest Is the Safety Boundary
How MCP fits behind a production adapter manifest with schemas, auth, approval modes, idempotency, observability, and replay.
Harness Improvement Loops Need Replayable Environments
Why harness improvement needs replayable episodes, bounded mutations, scorecards, source closure, and promotion gates.
AI Does Not Launch Once: Feedback Loops After Go-Live
A plain-English guide to operating agents after launch: corrections, recurring failures, proposal queues, rollout, rollback, and review.
How to Judge AI Work: Scorecards, Not Vibes
A practical guide for business teams evaluating AI agents with scorecards, examples, traces, human corrections, and launch gates instead of demos and vibes.
Trusting AI at Work: Approvals, Boundaries, and Receipts
A plain-English guide to agent trust: what AI can read, draft, send, change, approve, and how receipts make decisions accountable.
Before Your Team Asks for an AI Agent, Map the Real Work
A practical guide for business teams mapping real work before building agents: actors, evidence, tools, decisions, risks, exceptions, and feedback loops.
AI Agents for Business Leaders: Build the Airport, Not Just the Plane
A practical executive playbook for agentic AI: define the work, evidence, authority, scorecards, approvals, security, observability, and improvement loop.
Operating Agent Products: Feedback, Rollout, and the Improvement Loop
A PM operating model for shipped agents: trace review, corrections, proposal queues, scorecards, rollout, and rollback.
Trust Is a Product Surface: Approval Modes and Human Control for Agentic Products
How PMs should design trust for real agentic products: approval modes, human roles, evidence snapshots, DecisionRecords, policy gates, and graceful failure.
Scorecards Before Screens: Evals and Launch Gates for PMs Building Agents
A PM guide to defining agent quality with datasets, trace reviews, scorecards, release gates, and business metrics before building the agent UI.
The Control Tower Pattern: How PMs Should Design Multi-Agent Products
A PM guide to splitting multi-agent systems into specialist lanes while keeping orchestration governed and inspectable.
From PRD to Intent Catalog: The PM Spec for Agentic Products
How PMs turn vague agent ideas into intent catalogs, task templates, authority models, DecisionRecords, and launch criteria.
Product Managers: How to Think About and Build Complex Agentic Systems
A practical PM guide to building agentic systems with workflow maps, intents, context packs, tools, records, evals, and rollout gates.
How to Develop an Agent with an Agent Harness, End to End
An end-to-end field guide for building agents as measurable harnesses: context, planning, tools, records, evals, rollout, and learning.
Autotune the Harness: Baking the Improvement Loop into ContextOS
How ContextOS treats autotune as a gated loop over traces, scorecards, replay sets, bounded candidates, approval, and rollout.
Dataset-First Agent Engineering: The Golden Sets Behind Reliable Agents
A practical guide to golden sets, task distributions, corrected runs, held-out releases, and production slices for agent engineering.
How Great AI Engineers Build Agents: Datasets, Scores, and Harnesses That Improve
Why strong AI engineers build datasets, scorecards, traces, and improvement loops instead of treating agents as prompts plus tools.
Harness Candidates Are Model Checkpoints: How to Improve Agents Without Silent Mutation
How to treat every prompt, retrieval, tool, policy, and evaluator change as a scored, reviewed, reversible harness candidate.
Scorecards Over Vibes: The Five Metrics That Keep Agents Honest
The five metrics that keep agents honest: policy, utility, latency, safety, and economics.
Trace Review Is the Agent Debugger: Grade the Path, Not Just the Answer
How trace review grades the path, not just the answer, by inspecting context, plans, tools, guardrails, critic verdicts, and corrections.
Agentic AI Systems Before and After ContextOS
A table-first guide to why agentic systems need bounded context, governed tools, typed decisions, replay, evaluation, and controlled improvement.
AGENTS.md Done Right: The Navigation File That Actually Helps Coding Agents
How to write AGENTS.md as a short, scoped, testable navigation file for coding agents instead of a bloated prompt dump.
The Agent Harness Audit: A Production Readiness Checklist for Governed AI Agents
A production readiness audit for agent harnesses: forty-four runtime controls grouped into eight evidence-backed outcomes.
Replay Harness in Code: Reproducing a DecisionRecord Byte-for-Byte
A TypeScript build-along for replay: input loading, hash-chain verification, canonical loop replay, and DecisionRecord diffing.
End-to-End Refund: How 12 Primitives Compose in One Production Run
A single refund run traced through 12 ContextOS primitives, from invokeAgent envelope to byte-equal replay.
Failure Playbooks: The Typed Verdict Map
How to replace generic retry loops with typed failure verdicts, compensations, escalation paths, and reversal-token checks.
Approval Gates in Code: The Destructive-Mode Handshake
A build-along for approval gates: frozen evidence, human signatures, gateway redemption, and replayable destructive-action handshakes.
Build the Tool Gateway: The Boundary That Actually Stops a Bad Action
A build-along for the Tool Gateway: adapter manifests, typed envelopes, resolver checks, dispatch, and destructive-action boundaries.
The Critic: verify, score, consolidate — in 80 Lines
A compact Critic implementation that verifies plans, scores outcomes, consolidates results, and records caveats.
The Five Planes of Agentic Operating Systems
A working decomposition for production agent systems: Intelligence, Context, Decision, Action, and Trust.
Promotion-Aware Memory: Capture, Review, Promote, Recall in Code
A build-along for agent memory: capture, review, promote, recall, contradiction checks, and governed memory writes.
Build the Context Pack Compiler: Eight Stages, Eight Files
A build-along for the Context Pack compiler: eight deterministic stages that turn runtime inputs into a typed compiled context.
Context Graphs: Decision Lineage as a System of Record
How hash-chained DecisionRecords turn execution-time context into a queryable lineage graph for why an agent acted.
From Operator Correction to Released StrategyRule: The Improvement Loop, Coded
How one operator correction becomes a reviewed, replayed, versioned StrategyRule that prevents repeat agent failures.
Pack Rollout in Five Stages: Shipping a Context Pack Without Blowing Up Production
A five-stage rollout model for Context Packs: shadow, internal, low-risk, monitored expansion, full release, and rollback.
Replay Is the Real Audit Log
Why "we have logs" is not an audit story, and what a hash-chained Decision Record plus canonical replay actually buys you when an incident hits.
Wiring the Five Evaluators: Policy, Utility, Latency, Safety, Cost
A build-along for wiring policy, utility, latency, safety, and cost evaluators into a release-gated scorecard.
Context Packs in Practice: From Spec to Run
A practical walkthrough of Context Packs: buckets, policy bundles, evaluation gates, lifecycle, and the compile pipeline.
Building a Reliability Reviewer Agent: 70 Lines Past the Compliance One
How to extend the reviewer pattern for reliability: timeouts, retries, idempotency, fallback behavior, and rollback declarations.
Building a Compliance Reviewer Agent in 60 Lines and a Golden Set
How to build a compliance reviewer agent with a typed verdict envelope, rubric, golden set, and change-control queue.
Approval-Mode Tiers: A Risk Taxonomy You Can Actually Ship
Why ad-hoc approval gates rot in production, and how five canonical risk tiers turn governance from a meeting into a contract.
Beyond Prompts: The Architecture of Trust for Agentic AI
Building a governed decision runtime across Intelligence, Context, Decision, Action, and Trust — with evaluator scoring, approval tiers, and replay-bound audit.
Prompt Injection Is a Boundary Problem, Not a Prompt Problem
Why "smarter prompts" don't defend against indirect prompt injection, and what changes when authority lives outside the model's view.
Context Engineering in Production
Why most agent failures are not model failures — they are context failures — and what changes when context becomes a versioned, testable, replayable contract.