AI agents · Evals · Reliability
The model is one layer. The system decides whether it works.
Source-backed guides to the engineering around the model: context, tools, harnesses, loops, and the evidence that an agent did what it claimed.
Newest analysis
Agent systems · Source-backed analysis
AgenticOps Needs More Than Autonomy. It Needs Proof.
Cisco and Omdia's new research shows why network operations is moving from advisory AI to governed action. The next step is not more autonomy for its own sake, but an evidence loop that lets every production change be explained, verified, and recovered safely.
Reading paths
Start anywhere. Finish something.
Every piece stands alone, but these three sequences give the library an order and a reason to keep going.
Path 01 · 3 pieces · 31 min
Understand the system
The moving parts, before you choose a framework or a model.
- 01The Agent System Stack
- 02AI Agents, Explained
- 03How AI Agents Use Tools
Path 02 · 4 pieces · 42 min
Design the runtime
How context, harnesses, loops and state shape the work an agent can do.
- 01Context Engineering
- 02Harness Engineering
- 03Loop Engineering
- 04Human Steering in Agent Systems
Path 03 · 3 pieces · 34 min
Prove it works
From a promising demo to repeatable evidence and safer operation.
- 01AI Agent Evals
- 02Continuous Agent Assurance
- 03Agent Evaluations Are Security Systems
The library
Guides and analysis, newest first.
Cisco and Omdia's new research shows why network operations is moving from advisory AI to governed action. The next step is not more autonomy for its own sake, but an evidence loop that lets every production change be explained, verified, and recovered safely.
02Agent Oversight Needs More Than a DashboardEvaluation & safety39mAnthropic's new frontier-lab measurements make agent oversight more visible. The important engineering lesson is that activity metrics, control metrics, and outcome metrics answer different questions—and none can replace the others.
03A Managed Agent Harness Does Not Remove Your Evidence ProblemEvaluation & safety39mOpenAI's Agents API manages the Codex harness for long-running work. That changes who operates the runtime—not the evidence an engineering team needs before it can trust an agent's result.
04The More Agent Work You Run, the More Human Steering MattersAgent systems410mRunning more AI agents changes the bottleneck. The hard work moves toward task design, intervention, verification, and integration—not away from people altogether.
05Agent Evaluations Are Security SystemsEvaluation & safety511mA realistic agent evaluation needs more than a benchmark and a score. It needs explicit scope, least-privileged tools, observable actions, stop conditions, and a recovery plan.
06Beyond the Prompt: The Agent System Stack ExplainedAgent systems912mContext, tools, memory, harnesses, loops, and evals are different layers of an AI agent system. Here is how they fit together—and why the model alone does not decide whether an agent works.
By subject
Three subjects with enough written to be worth a page.
7 pieces
Agent systems
What an agent is made of, and which layer is actually deciding the outcome.
3 pieces
Context, memory and tools
What the agent can see, what it retains, and what it is able to act on.
5 pieces
Evaluation and safety
How you find out whether it worked, and how you keep knowing after it ships.
The Runnable Weekly
A source-linked guide to the week worth a builder’s attention.
Issue 07 · September 20, 2026
September 8: Evidence Becomes Operational
New frontier-lab measurement proposals, an embedded-evaluation partnership, and managed-agent tracing all point to the same practical requirement: a serious agent system needs evidence that connects authority, action, verification, and recovery.
Issue 06 · September 8, 2026
September 1: Agent Capability Changes Operations
GPT-6 Astra’s restricted deployment, an independently documented public-agent coordination incident, and Claude Fable 5.1’s long-run economics all point to the same shift: capable agents change the operating system around the model.
Issue 05 · September 7, 2026
August 24: Faster AI and Autonomous Weapons
Gemini 3.7 Flash pushed harder on price and speed, OpenAI published its first Jalapeño chip results, Qwen brought a frontier-scale model into the open, and autonomous targeting moved from theory to reported real-world harm.
Past weeks stay here as a dated archive. Nothing is emailed and nothing is gated.