OpenAI’s September announcements change the economics of delegation. The interesting question is whether a cheaper model and an always-running agent reduce the total work required to reach an acceptable result. That includes the user’s review, correction, and recovery effort.
This assessment extends our Astra launch analysis. It focuses on the September 29 releases, with evidence checked through October 6, 2026. Vendor measurements remain vendor measurements; the evaluation designs below are proposals.
What changed after Astra
OpenAI introduced GPT-6.1 Sol on September 29. Its launch page lists standard prices of $2 per million input tokens, $0.10 for cached input, and $10 for output. It reports results approaching Astra on several coding and professional-work evaluations, while retaining Astra’s advantage on difficult scientific work. The page also explains that its factuality sample consists of conversations selected because users had flagged prior errors. It is not representative of ordinary traffic. GPT-6.1 Sol announcement
DevDay also introduced Dots for continuing responsibilities, added computer use to the Agents API, and announced event-triggered plugin automations. Availability differs: Dots has plan and market restrictions, the Decisions API was a limited preview, and Private Inference was described as a forthcoming preview. These are separate release states, not one universally available platform bundle. DevDay recap
| Development | Evidence supports | Assessment question |
|---|---|---|
| Sol pricing and launch evaluations | A new candidate for cost-sensitive professional work | Does it preserve acceptance rates on our difficult cases? |
| Persistent Dots | A product for ongoing responsibilities | Can ownership, scope, and permissions survive personnel changes? |
| Hosted computer use | A managed execution surface | Can the application reconstruct what changed in external systems? |
| Event-triggered automations | Work can start from connected-app events | Can replayed or forged events trigger duplicate effects? |
The relevant cost denominator is accepted work
For a proposed evaluation, define cost per accepted outcome as total inference, tool, infrastructure, and review cost divided by the number of independently accepted outcomes. Failed attempts and abandoned tasks belong in the numerator. A model that writes cheaper drafts can still increase this measure if its mistakes are expensive to detect.
Caching adds another variable. A workflow with a stable reference corpus differs from one whose first messages change every turn. Measure cache behavior using the actual production prompt assembly order. Do not extrapolate a cached-input price to all input tokens or assume that a long context guarantees reuse.
A useful experiment has two passes. First, keep the harness, tools, task set, retry budget, and acceptance rubric fixed while comparing model configurations. Second, allow each configuration a documented tuning budget. The first estimates a substitution effect; the second estimates the value achievable after integration work. Mixing the passes makes a model improvement indistinguishable from a better harness.
Report results by task class. Routine extraction, ambiguous investigation, and code migration should not share a single acceptance threshold. A pooled average can hide a regression in the small set of tasks where mistakes have the greatest consequences.
Persistent agents turn event handling into a product requirement
Consider an illustrative account-research agent that prepares a briefing whenever a sales opportunity changes. A webhook is retried, two colleagues edit the opportunity, and the owner leaves the company before the draft completes. The language task is straightforward. The lifecycle is not.
The application needs a stable identity for the triggering event, a version of the objective, a current owner, and a clear distinction between preparing a draft and delivering it. Before an external write, it must check current authority and current destination state. A remembered instruction can explain why work started; it cannot establish that the same person still has permission to finish it.
This is our engineering inference from the move toward continuing responsibilities. It is not a claim that Dots lacks these controls or that the launch announcement proves them. A procurement review should request evidence about revocation, retained credentials, event deduplication, and exportable execution records for the chosen deployment.
Computer use makes partial completion consequential
A browser workflow may save a record before the agent sees a timeout. Retrying the whole task can create a duplicate. Conversely, a confident completion message can arrive while a submission remains pending.
Evaluate the external state separately from the conversational answer. Use a controlled application with authoritative operation identifiers and a way to inspect postconditions. Introduce an interruption after submission, then resume the session. The desirable result is a reconciled operation with a trace explaining what happened. A second successful submission is a failure if it duplicates the first.
The new computer-use surface is documented in the September 29 API changelog. That entry establishes product support, not exactly-once semantics for arbitrary websites. OpenAI API changelog
A pilot that can change the adoption decision
Use a held-out collection of real, consented work with frozen input snapshots. Include successful ordinary cases, ambiguous requests, denied actions, unavailable tools, and interrupted writes. Randomize assignment and blind artifact reviewers to the model. Predeclare acceptable quality loss and maximum review burden rather than choosing thresholds after seeing results.
Track acceptance rate, cost per accepted outcome, median and tail completion time, review minutes, unauthorized attempts, and unresolved external effects. Count a safe refusal separately from a capability failure when the task requests an impermissible action. Count a refusal on legitimate work as a usability cost.
My assessment: Sol deserves evaluation as an economical default, while persistent agents deserve evaluation as a new operational system. Neither decision follows from the other. Cheaper inference can finance more verification, but it can also finance more unchecked activity; the application determines which occurs.