AI agents · Evals · Reliability

The model is one layer. The system decides whether it works.

Source-backed guides to the engineering around the model: context, tools, harnesses, loops, and the evidence that an agent did what it claimed.

GOALACTOBSERVECHECKbounded · observablestoppable · recorded
act → observe → adjust → stop

Newest analysis

The full library →

Agent systems · Source-backed analysis

AgenticOps Needs More Than Autonomy. It Needs Proof.

Cisco and Omdia's new research shows why network operations is moving from advisory AI to governed action. The next step is not more autonomy for its own sake, but an evidence loop that lets every production change be explained, verified, and recovered safely.

PublishedSeptember 29, 2026
Reading9 min
Sections6
Primary sources4

Reading paths

Start anywhere. Finish something.

Every piece stands alone, but these three sequences give the library an order and a reason to keep going.

The library

Guides and analysis, newest first.

Filter all 15 →
01AgenticOps Needs More Than Autonomy. It Needs Proof.Agent systems49m

Cisco and Omdia's new research shows why network operations is moving from advisory AI to governed action. The next step is not more autonomy for its own sake, but an evidence loop that lets every production change be explained, verified, and recovered safely.

Agent systems9 min read4 sources
02Agent Oversight Needs More Than a DashboardEvaluation & safety39m

Anthropic's new frontier-lab measurements make agent oversight more visible. The important engineering lesson is that activity metrics, control metrics, and outcome metrics answer different questions—and none can replace the others.

Evaluation & safety9 min read3 sources
03A Managed Agent Harness Does Not Remove Your Evidence ProblemEvaluation & safety39m

OpenAI's Agents API manages the Codex harness for long-running work. That changes who operates the runtime—not the evidence an engineering team needs before it can trust an agent's result.

Evaluation & safety9 min read3 sources
04The More Agent Work You Run, the More Human Steering MattersAgent systems410m

Running more AI agents changes the bottleneck. The hard work moves toward task design, intervention, verification, and integration—not away from people altogether.

Agent systems10 min read4 sources
05Agent Evaluations Are Security SystemsEvaluation & safety511m

A realistic agent evaluation needs more than a benchmark and a score. It needs explicit scope, least-privileged tools, observable actions, stop conditions, and a recovery plan.

Evaluation & safety11 min read5 sources
06Beyond the Prompt: The Agent System Stack ExplainedAgent systems912m

Context, tools, memory, harnesses, loops, and evals are different layers of an AI agent system. Here is how they fit together—and why the model alone does not decide whether an agent works.

Agent systems12 min read9 sources