The Runnable Weekly · The Runnable Weekly · Issue 05
Week of August 24: Faster Models, Custom Chips, and Autonomous Weapons
This week’s important AI stories happened at very different layers of the stack. Google made a capable agent model cheaper to run. OpenAI showed why it is designing its own inference hardware. Qwen widened access to open models built for long-running work. And reporting from Ukraine showed why the question of who controls an autonomous system stops being abstract when software can select a target in the physical world.
What to remember
- A faster or cheaper model matters when it reduces the cost of completing a whole task—not merely the cost of generating one token.
- Agent latency compounds across planning, tool calls, retries, and verification, which makes inference hardware part of the agent experience.
- Open weights give teams more control over deployment and evaluation, but published benchmarks still need to be reproduced on the work that matters to them.
- When an AI system can affect the physical world, authority, human review, and accountability must be designed before deployment rather than inferred after an incident.
The short version
- Google introduced Gemini 3.7 Flash with higher reported coding and agent performance than 3.6 Flash, while temporarily pricing it at half the earlier model’s original per-token rate.
- OpenAI published the first performance results for Jalapeño, its custom inference chip, arguing that lower latency and better performance per watt matter especially for multi-step agents.
- Qwen released open weights for Qwen3.8 models aimed at coding, research, and long-horizon agent tasks, giving builders another model family they can inspect and run in their own environments.
- A New York Times investigation reported that an experimental AI-guided Russian drone selected its final target at a gas station in Ukraine, where its explosion killed three civilians. The account turns human control from a design preference into an accountability question.
The common thread is not that AI became smarter in one dramatic jump. The stack became easier to scale: cheaper models, specialized chips, and more deployable open weights. At the same time, the drone report is a reminder that scaling autonomy without equally clear control can make the consequences harder to contain and harder to assign.
1. The useful model race is moving toward work per dollar
- Introducing Gemini 3.7 Flash — Google
- Gemini 3.7 Flash model card — Google DeepMind
Google released Gemini 3.7 Flash on August 13, only three weeks after 3.6 Flash. Google reports gains over 3.6 Flash on production-code, long-horizon software-engineering, web-development, and workflow-automation evaluations. It also introduced temporary pricing of $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Those are vendor-reported results and introductory prices, so they should be treated as a reason to test the model rather than proof that every application will improve.[1][2]
For agent builders, the meaningful measurement is the cost of a successful run. A model can be inexpensive per token and still be costly if it needs more retries, produces longer traces, or requires more human correction. Compare total task cost, elapsed time, intervention count, and verified success on your own workflows before changing a production route.
2. Inference hardware is becoming part of agent design
On August 25, OpenAI published its first measured results for Jalapeño, the company’s custom inference chip. Across three public models in OpenAI’s tests, the company reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. The tests were conducted and reported by OpenAI, even though they used the public InferenceX benchmark, so independent reproduction will matter.[3]
The system-level point is easy to miss. An agent may wait for the model dozens of times while it plans, calls tools, reads results, retries, and verifies its work. Small delays compound across that loop. Hardware that improves both latency and throughput can change how responsive an agent feels and how many concurrent runs an organization can afford—even if the model itself does not change.
3. Open models are competing for agent workloads
Qwen released its Qwen3.8 open models in mid-August, including a large mixture-of-experts model and a 27-billion-parameter dense model. The official repository emphasizes coding, research, long-horizon agent execution, adjustable reasoning effort, and compatibility with common serving frameworks. It also provides instructions for serving the models through OpenAI-compatible APIs.[4]
The important part is deployment choice. Open weights let a team inspect versions, control where inference runs, and test the model inside its own harness. They do not remove operational work: tool-call formats, memory limits, hardware requirements, licensing, and real task performance still have to be checked. The strongest reason to try an open model is control over the experiment, not a leaderboard claim.
4. Autonomy has consequences beyond the interface
- A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. — The New York Times
A New York Times investigation published on August 24 reported that an experimental AI-guided Russian drone killed three civilians during a July 6 attack in Zaporizhzhia, Ukraine. According to the experts, military officials, and forensic team cited by the Times, human operators sent the drone toward a gas station, while its onboard system selected the exact target—likely propane tanks—without a human pilot making that final choice. The report describes this as the first documented civilian deaths attributed to this kind of Russian autonomous targeting, not as the first use of automation in warfare generally.[5]
That distinction matters. A human still chose the mission and destination, but software reportedly chose the final aim point. When an agent moves from recommending an action to selecting and executing one, the control problem changes: who can stop it, which uncertainty should force a handoff, what evidence is recorded, and who is accountable for a mistake? Those questions apply far beyond weapons, but here the cost of getting them wrong is irreversible.
Primary references
Sources
These references support the definitions and technical claims in this article. Product-specific guidance is identified by its publisher.
- 01Introducing Gemini 3.7 FlashGoogle ↗
- 02Gemini 3.7 Flash model cardGoogle DeepMind ↗
- 03Jalapeño’s first results show speed and efficiency in AI inferenceOpenAI ↗
- 04Qwen3.8 official repository and model informationQwen Team ↗
- 05A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.The New York Times ↗