The Runnable Weekly · The Runnable Weekly · Issue 07
Week of September 8: Evidence Becomes an Agent-Operations Requirement
This week’s stories appear to be about very different things: a proposal for measuring AI development inside frontier labs, an arrangement to place evaluators inside one of those labs, and the tracing facilities of a managed-agent API. For builders, they converge on one question. When an agent system matters, can someone reconstruct what authority it had, what it did, how the outcome was checked, and what happened when the system needed correction?
What to remember
- A published activity metric is useful only when its definition, limits, and connection to a decision are inspectable.
- An embedded evaluator is not a safety certificate. Independence depends on concrete access, reporting, funding, and escalation arrangements that can be examined.
- A trace records activity; an operating record connects that activity to authority, verification, and recovery.
- The smallest practical improvement is to preserve one reviewable record per consequential agent task, rather than relying on a productivity dashboard or final chat message.
The short version
- On September 17, Anthropic proposed recurring measurements intended to make frontier-lab development more visible. Its proposal explicitly warns that best-effort classifications, short snapshots, and compute shares have important limits.
- On September 18, Anthropic announced a planned embedded-evaluation partnership with Accenture. The companies describe evaluator access comparable to an employee’s, but Anthropic also says standards for access, reporting, and funding are not yet settled.
- OpenAI’s Agents API documentation describes traces that can expose a session’s inputs, outputs, duration, status, tool activity, and delegated work, and that can be exported through an API. Those capabilities make investigation easier; they do not by themselves establish that an outcome was correct or that authority was justified.
The shared engineering lesson is that agent oversight cannot end with a number, a partner announcement, or a log stream. Each is an input to an operating record. That record should let a reviewer move from the job the system was asked to do, through the authority and material actions it used, to the external check that established success or failure—and to the recovery decision when the check did not pass.
1. A metric matters when it can change a decision
Anthropic’s September 17 proposal is valuable precisely because it is cautious about what its measurements can establish. The framework discusses AI participation in research and development, agent supervision, and compute allocation, while warning that its classifications are best effort, that a one-week snapshot is not enough to establish a trend, and that compute share is not a direct measure of safety-work quality. The result is a more useful model for operational metrics: preserve the method, state the uncertainty, and do not let a convenient number stand in for a broader conclusion.[2]
For a team operating agents, the equivalent is to link every important metric to a concrete decision. A rising intervention rate might trigger a review of task scope or instructions. A failed verification might block a state change. A growing retry count might cap concurrency or require a human handoff. Without that link, the number is mainly a report about activity. With it, the number becomes part of a control system.
2. Independent evaluation is an operating design, not a label
- Partnering with Accenture on embedded evaluation — Anthropic
Anthropic and Accenture announced a partnership for independent evaluation of frontier AI, led by Accenture’s Faculty business. Anthropic says the planned work includes evaluation and red-teaming, alignment assessment, and testing of model safeguards. It describes embedded evaluators as working inside the company with access comparable to an employee’s, but also says that shared standards for information access, reporting, and funding do not yet exist. The announcement is a commitment to build an evaluation arrangement, not a published evaluation result or safety certification.[3]
That distinction is practical for any organization. If a reviewer is expected to make an agent system more trustworthy, define the reviewer’s scope before the system needs one: which runs and artifacts can be inspected; which identities, policy decisions, and failures are visible; who may pause work; how disagreements are recorded; and what can be disclosed. An evaluator who can see a final dashboard but not the task definition, permissions, raw evidence, or remediation history cannot independently establish much.
3. Tracing is the beginning of evidence, not the end
- Tracing — Agents API — OpenAI Developers
- Create an agent session — OpenAI Developers
OpenAI’s Agents API documentation describes session traces that record steps in a turn, including model responses, tool calls, and delegated subagent work. The trace view includes recorded inputs, outputs, duration, and status, and the documentation describes exporting session traces in OpenTelemetry format. Those details are important because a long-running agent cannot be investigated from a final answer alone. A trace gives an operator a path back through the work.[4][5]
But a trace can still be incomplete evidence. It can show that a tool call happened without showing whether the tool should have been available, whether its result was independently checked, or whether a failed action left the environment in a recoverable state. Treat the trace as one layer in a task record, alongside the task contract, resolved permissions, environment version, material artifacts, acceptance check, and any recovery or escalation decision.
Task
What outcome was requested, and what external check defines success?
Authority
Which identity, tools, data, and approval boundaries were available?
Trace
Which material actions, tool results, subagents, retries, and interventions occurred?
Evidence
Which artifact or environment state did an independent check accept or reject?
Recovery
If the check failed or work stopped, what state remained and who owned the next decision?4. One small test for this week
- Choose one bounded workflowPick an agent-assisted task with a real external outcome: a code change with tests, a support draft with an approval step, or a data update with a reconciliation check.
- Write the task and authority contract firstState the desired outcome, the allowed tools and data, the person who can approve a consequential action, and the exact condition that should stop the run.
- Keep the execution recordSave the resolved configuration, tool and subagent trace, material artifacts, and the actual final environment state. Do not rely on a prose completion message.
- Verify outside the agentRun a deterministic check, compare a source record, or ask a designated reviewer to decide whether the result is acceptable. Record the result next to the trace.
- Simulate one imperfect conditionDeny an unnecessary permission, inject a stale input, or make the verification fail. Confirm that the system pauses or leaves enough evidence for someone to recover safely.
This is not a test of whether the model can sound convincing. It is a test of whether the operating system around the model remains legible when the work is incomplete, surprising, or wrong. That is the standard a metric, evaluator, and trace should ultimately support.
Primary references
Sources
These references support the definitions and technical claims in this article. Product-specific guidance is identified by its publisher.
- 01Last Week in AI #344 — Navier–Stokes, Pacing the Frontier, AI MisuseLast Week in AI ↗
- 02Measurements for understanding the pace of AI development inside frontier labsAnthropic ↗
- 03Partnering with Accenture on embedded evaluationAnthropic ↗
- 04Tracing — Agents APIOpenAI Developers ↗
- 05Create an agent sessionOpenAI Developers ↗