Evaluation & safety · Source-backed analysis
A Managed Agent Harness Does Not Remove Your Evidence Problem
A managed agent runtime can remove real engineering work: context compaction, tool orchestration, sandbox provisioning, durable sessions, and subagent coordination. It does not remove the obligation to prove what the agent did. The useful question is not whether the runtime is managed. It is whether a team can reconstruct the work, verify the result, and recover safely when a run does not finish as planned.
What to remember
- A managed harness owns execution infrastructure; the application team still owns task definition, authority boundaries, acceptance criteria, and the decision to trust a result.
- Agent completion is a claim. Tests, resulting state, action records, and reviewable artifacts are the evidence that supports or rejects it.
- Durable sessions improve continuity, but continuity is not recoverability unless the next operator can see the relevant state and make a safe decision.
- The smallest credible operating record includes the task, authority, environment, action trail, resulting artifacts, verification results, and recovery decision.
What a managed harness actually changes
- Introducing the Agents API — OpenAI
OpenAI introduced the Agents API in public beta on September 10, 2026 as managed infrastructure for cloud agents built on the Codex harness. Its role is operational: it coordinates model calls, context, tools, sessions, subagents, and execution environments. Developers supply the task, model, tools, and environment; OpenAI operates and maintains the harness.[1]
That is a meaningful division of labor. A team no longer has to implement every detail of a long-running loop before it can use one. The API can compact earlier context as sessions approach their context limit, select tools on demand, coordinate parallel subagents, and run work in an OpenAI-hosted sandbox, its own infrastructure, or a partner environment.[1]
- The runtime can manage context; the team still decides which requirements and constraints are authoritative.
- The runtime can invoke tools; the team still decides which tools exist, what they are permitted to change, and when approval is required.
- The runtime can preserve a session; the team still decides what counts as a completed, correct, and safe outcome.
- The runtime can return artifacts; the team still needs an independent way to inspect and verify those artifacts.
Completion is a claim; evidence is the answer
- Agents API reference — OpenAI Developers
An agent's final message is useful as a summary. It is not proof that a task succeeded. A coding agent can say that it fixed a bug while tests still fail. A research agent can say that it found the answer while citing the wrong source. An operations agent can say that it mitigated an incident while the service state remains unhealthy. The environment—not the prose—decides whether the claim holds.
The Agents API exposes several surfaces that make a stronger record possible: session turns, session items, subagent records, event streams, and artifacts. In an OpenAI-hosted session, an artifact is an immutable file published by a completed turn. Those primitives make a reviewable operational trail feasible. They do not automatically supply an acceptance test or decide whether the outcome is good enough for the application.[2]
Task
Versioned request, explicit acceptance criteria, and a named owner.
Authority
Tools, credentials, approval rules, and the actions the agent may not take.
Execution record
Environment version, input state, material tool actions, observations, and handoffs.
Result
Reviewable diff, report, created artifact, or resulting system state.
Verification
Tests, policy checks, independent queries, or human review that directly check the result.
Recovery decision
Whether retry is safe, what state persists, what must be rolled back, and who decides next.Durable sessions are valuable. Recoverability is a separate property.
- Introducing the Agents API — OpenAI
- Agents API reference — OpenAI Developers
- Degraded Performance affecting Agents API — resolved September 14, 2026 — OpenAI Status
Long-running work needs continuity. The Agents API is designed to carry work across multiple context windows, and the API exposes session, turn, event, item, subagent, and artifact resources. This gives a system a durable place to keep work in progress instead of compressing the entire task into one response.[1][2]
But a session that persists is not automatically a task that can be safely resumed. Recovery requires an operator to know what happened, what remains uncertain, and which actions are safe to repeat. That distinction becomes practical during ordinary service failures. On September 14, OpenAI reported an incident in which some managed Agents API sessions experienced delays or could not start turns; the incident was later marked resolved. The important engineering lesson is not to infer a cause from that event. It is to design every agent workflow so an interruption produces a clear decision point rather than an ambiguous restart.[3]
- Make side effects idempotent where possibleA retry should not silently create a second ticket, send a duplicate message, or apply a change twice. Store an operation identifier or an equivalent external receipt before the agent continues.
- Separate progress from successRecord the latest completed checkpoint and the evidence for it. Do not represent a partially completed task as a completed outcome merely because a session still exists.
- Require a verification step after recoveryWhen a task resumes, re-check the resulting environment before taking the next irreversible action. State can persist while the external world has changed.
- Keep a human escalation pathAn uncertain result, missing artifact, failed verification, or expanded authority boundary should stop automation and route the decision to a named person.
The operating standard: outsource orchestration, not accountability
The practical value of a managed harness is that it lets a team spend less time rebuilding durable agent infrastructure and more time on the parts that make the application specific: the workflow, domain knowledge, authority model, and evaluation. That is the right trade. The mistake is to treat the abstraction boundary as an accountability boundary.
Before a managed agent is allowed to perform meaningful work, the team should be able to answer six questions without hesitation: What exactly was it asked to do? What could it access or change? What did it actually do? What did it produce? How was the result checked? What happens if the run stops halfway through? A system with firm answers to those questions is inspectable. A system that relies on a confident final message is merely persuasive.
- Start with bounded tasks whose final state can be checked independently.
- Keep destructive or externally visible actions behind explicit approval boundaries.
- Store the smallest complete evidence package with the result, not in a separate oral history of the run.
- Treat every interruption, failed check, and human intervention as input to a better task contract or harness configuration.
- Judge the workflow by verified outcomes and safe recovery, not by how autonomous it looked from the outside.
Primary references
Sources
These references support the definitions and technical claims in this article. Product-specific guidance is identified by its publisher.