The phrase “AI agent” hides three different products: a personal agent that helps one person finish work, an execution agent that operates on code and systems, and a deployed service agent that acts for a company under policy.
OpenAI now has a named product at each layer. ChatGPT Work is the general work surface. Codex is the execution substrate embedded across technical and increasingly non-technical workflows. OpenAI Presence is a deployed enterprise product for specific voice and chat jobs.
They share technology, but they are not interchangeable. Understanding the boundaries is more useful than treating all three as “ChatGPT with tools.”
The three layers
| Product | Primary user | Unit of work | Control model | Availability signal |
|---|---|---|---|---|
| ChatGPT Work | Individual knowledge worker or team | A project that may span apps, files, and hours | User steers and reviews through ChatGPT | Rolling out across ChatGPT plans and desktop surfaces |
| Codex | Developer, technical professional, and tool-building knowledge worker | A task executed against repositories, terminals, files, and connected systems | Workspace instructions, sandboxing, approvals, review, and version control | Broad product availability; integrated into the ChatGPT desktop app |
| OpenAI Presence | Enterprise operations team | A defined customer or internal workflow | Policies, allowed actions, simulations, guardrails, escalation, and controlled improvement | Limited general availability for eligible enterprise customers |
The availability descriptions above come from the July product announcements and can change. Check current release notes before making a purchasing decision.
ChatGPT Work: from conversation to deliverable
OpenAI describes ChatGPT Work as an agent that can gather information across apps and workflows, create finished sheets, slides, documents, and web apps, and continue complex projects for hours. GPT-5.6 powers the surface, with Codex technology handling execution.
The July announcement also says the Codex app is merging into the ChatGPT desktop app while retaining Codex as a distinct coding and technical work mode. The desktop integration adds a built-in browser, local files and apps, multi-repository projects, inline diff editing, and pull-request review.
The product implication is larger than a navigation change. Chat, document production, browser work, and code execution now share one project surface. That reduces handoffs between “ask,” “make,” and “verify.” It also increases the importance of project-level context, permissions, and review because the same interface can move from reading information to changing an artifact.
Codex: the execution layer is escaping the engineering department
Codex began as a coding agent, but OpenAI increasingly frames it as a general execution environment. Its study of Codex usage reports four trends:
- longer delegated task horizons;
- Codex becoming the primary AI work tool across departments inside OpenAI;
- faster growth among non-developer users;
- more people using it for work outside their formal job category.
The reported numbers are striking, but they need the right interpretation. Task duration is model-estimated, not observed human time saved. Internal OpenAI adoption is an extreme early-adopter setting. Token share is not the same as productivity. The evidence supports a behavioral claim—people are delegating broader and longer work—not a universal ROI claim.
OpenAI’s later task-crossover research strengthens the same direction. In a sample of more than 800,000 U.S. work-related ChatGPT messages, OpenAI classified 43.5% of occupation-specific messages as tasks associated with another occupation. The likely organizational effect is fewer small handoffs: a marketer troubleshoots a site, a salesperson explores data, or a small-business owner drafts and analyzes without waiting for a specialist.
That can increase autonomy and speed. It can also move work outside the review structures that previously came with the specialist. When capability crosses a job boundary, governance has to cross with it.
Presence: OpenAI packages the missing production layers
Presence is the most revealing product in the set because its announcement explicitly says a model is insufficient.
A Presence deployment begins with a specific job. The agent receives only the knowledge and system access required for that job. The company defines what it can do, when approval is required, and when a person takes over. Simulations and graders test common, edge, and higher-risk scenarios before launch. Production sessions and escalations become evidence for improvements, while Codex proposes changes that teams test and approve before rollout.
That is an agent operating model:
- Define a job and success criteria.
- Limit context and capabilities.
- Encode policy and escalation.
- Simulate before launch.
- Observe production sessions.
- Propose a versioned change.
- Test it against the current production version.
- Approve a controlled rollout.
OpenAI reports that Presence powers its English-language phone support and resolves 75% of inbound issues without human assistance, with a 15-point reduction in handoffs during one ten-day improvement period. Those are OpenAI’s own operational claims for a particular deployment. They are useful evidence that the system runs in production, not a forecast for another company’s workflow.
A maturity ladder for adopting the stack
Do not start with the most autonomous surface. Start with the smallest unit that produces evidence.
Level 1: assisted work
The person remains in the loop for every output. The agent drafts, analyzes, or edits; the person decides and acts.
Exit evidence: repeated useful output on a defined task, with known failure modes.
Level 2: delegated artifact production
The agent creates a complete file, analysis, code change, or project for review.
Exit evidence: acceptance rate, correction categories, reproducible validation, and bounded tool access.
Level 3: supervised action
The agent reads from systems and proposes or performs reversible actions with approval.
Exit evidence: policy adherence, correct approvals, idempotency, rollback, and trace completeness.
Level 4: scoped production operation
The agent handles a defined workflow, escalates exceptions, and is monitored after launch.
Exit evidence: outcome quality, escalation quality, incident rate, drift detection, and a safe improvement process.
Presence is positioned at Level 4. ChatGPT Work and Codex can participate at every level, but the product name does not grant the surrounding controls automatically.
The enterprise design questions that remain
The OpenAI stack makes execution easier. Buyers still own several decisions:
- Who is accountable for the completed outcome?
- Which identity does the agent use for every external action?
- What data is allowed into the task context?
- Which actions require approval, and from whom?
- How is a partially completed multi-system task reversed?
- Which production sessions become evaluation cases?
- Who can approve a policy or prompt change?
- What happens when the model, connector, or price changes?
OpenAI Presence answers some of these through a deployed engagement. A team building with ChatGPT Work, Codex, or the API needs to make the answers explicit in its own operating model.
The durable takeaway
OpenAI’s product strategy is converging on delegated work, but delegation comes in layers. ChatGPT Work is a general user agent. Codex is an execution environment. Presence is a governed enterprise deployment product. The distinction determines who supplies the policy, evaluation, approval, and operations layer.
The buying question is therefore not “Which agent is smartest?” It is “Which layer are we adopting, and which production responsibilities remain ours?”