A weekly field guide
The Runnable Weekly
The few developments worth a builder’s attention, separated into what happened, why it matters in practice, and the original reporting for anyone who wants to go deeper. 4 issues so far, 12 primary sources between them.
Latest · Issue 04
July 27: Safety Boundaries and Evidence
The week’s safety stories were not abstract: they concerned agents finding unintended routes through environments and models exploiting the rules of an evaluation. The practical response is better boundaries and better evidence.
Earlier weeks
Issue 03 · May 27, 2026
May 18: From Demos to Verifiable Work
Google’s agent-first developer tooling, a real-world terminal benchmark, and a model-assisted mathematical result all made the same case: autonomy only becomes useful when there is an environment and a credible way to check the result.
Issue 02 · May 5, 2026
April 27: Long Context and Agent Tests
A model with longer context, a benchmark for changing work environments, and safety research on delegated work all pointed to the same thing: an agent must handle state and incentives, not just answer a prompt.
Issue 01 · March 23, 2026
March 16: Agent Runtimes and Better Checks
Three developments from the week of March 16: agent runtimes started to become a product category, generative models moved deeper into interactive software, and model releases made evaluation more important—not less.