Agentic AI
21 essays tagged with Agentic AI.
Agentic Recommender Systems: How Harness Engineering Changes Personalization
Recommenders are evolving from rankers into tool-using agents. Here is the hybrid architecture, evaluation harness, and rollout plan needed to build them safely.
Agent Evaluation After the Leaderboard: The KDD 2026 Production Assurance Playbook
KDD 2026 shows why agent evaluation must cover protocols, trajectories, external state, evaluator operating profiles, causal rollout, and drift.
Anthropic's Latest Work: Safety Moves Into the Runtime
A source-backed field report on Fable 5's biology safeguard update, enterprise inference hooks, skill and plugin scanning, Claude Code security, and Anthropic's global policy expansion.
OpenAI's Latest Work: A 10-Day Field Report
A source-backed 10-day delta across Codex, plugins, API operations, model economics, security scanning, and the GPT-5.4 migration.
Anthropic in 2026: A Full-Stack Map of Its Latest Work
A source-backed map of Anthropic's latest work across Claude models, agents, the developer platform, safety, research, enterprise strategy, and compute.
Anthropic's 2026 Safety Stack: Alignment, Safeguards, Containment, and Disclosure
A critical map of Anthropic's latest safety work across constitutional training, capability evaluations, classifiers, model routing, containment, trusted access, and incident disclosure.
From Claude Code to Cowork and Claude Tag: Anthropic's Agent Product Stack
How Claude evolved from chat to coding, cross-app work, and a persistent team agent—and what identity, memory, permissions, and oversight each surface requires.
ChatGPT Work, Codex, and OpenAI Presence: The Emerging Agent Stack
How ChatGPT Work, Codex, and OpenAI Presence form distinct layers of OpenAI's shift from helpful answers to governed, long-running work.
GPT-5.6 Field Guide: Models, Tools, Reasoning, and the New Runtime
A practical guide to GPT-5.6 Sol, Terra, and Luna—and the runtime features that matter more than a headline benchmark score.
OpenAI in 2026: A Full-Stack Map of Its Latest Work
A source-backed map of OpenAI's newest work across models, agents, multimodal products, science, safety, and compute—and how the pieces fit.
Inside OpenAI's 2026 Agent Platform: Responses, Tools, MCP, and Control
A builder's map of the Responses API, programmatic tools, multi-agent orchestration, MCP, identity, spend controls, and deployment boundaries.
Give Claude Code, Cursor, and Codex Persistent, Auditable Memory
Coding agents are brilliant and amnesiac. SecondBrain's open-source Memory API gives Claude Code, Cursor, Codex, and ChatGPT shared, local-first memory over HTTP and MCP — where every answer carries a citation back to the source chunk.
SecondBrain: A Local-First Agent Operating System You Can Run, Inspect, and Trust
SecondBrain is an open-source, local-first agent OS: cognition, memory, governed tools, durable sessions, workflows, and bounded self-improvement in one inspectable runtime you run on your own machine. Here is how it works and how to run it.
The State of AI Agents in 2026: Standards Converged, Models Improved, Production Moved to the Harness
A mid-2026 review of agentic AI: MCP, A2A and AP2 converged as standards and models got more reliable — yet the bottleneck moved to the governed agent harness.
AI Tokenomics: From Cost per Token to Cost per Trusted Outcome
AI tokenomics connects cost per token, agentic cost multipliers, routing, evals, governance, and cost per trusted outcome.
The Autonomy Budget: How Enterprises Should Decide What AI Agents Are Allowed to Do
A practical governance model for granting AI agents bounded authority based on risk, evidence, policy confidence, evals, and approval.
Antahkarana Stack: A Cognitive Layer for Local-First Agents
A builder-facing explanation of Antahkarana as an engineering layer inspired by the inner faculties of Manas, Buddhi, Chitta, and Ahamkara.
ContextOS: A Research-Grounded Architecture for Governed Agent Runtimes
A research-grounded framing of ContextOS as a governed runtime for context, tools, memory, security, evaluation, replay, and optimization.
AI Agents for Business Leaders: Build the Airport, Not Just the Plane
A practical executive playbook for agentic AI: define the work, evidence, authority, scorecards, approvals, security, observability, and improvement loop.
Agentic AI Systems Before and After ContextOS
A table-first guide to why agentic systems need bounded context, governed tools, typed decisions, replay, evaluation, and controlled improvement.
The Five Planes of Agentic Operating Systems
A working decomposition for production agent systems: Intelligence, Context, Decision, Action, and Trust.