Issue 4 October 2026
-
The two-dollar frontier
Google’s Gemini 4 Argon is impressive. More striking is that top-tier intelligence now has a going rate
Read the leader →Google’s new frontier model, Gemini 4 Argon, comes with the usual flourish of benchmarks. It is state of the art on DeepSWE v1.1, a test of long, real-world software engineering, at 77.9%. It tops the Vals Index of economically weighted knowledge work, leads Zapier’s AutomationBench at 51.3% and scores 91.7% on LVBench, a long-video test. Its maximum output grows from 64,000 tokens to 1m, enough for a single chain of reasoning to run to hundreds of thousands of tokens. The internal anecdotes are more persuasive than the leaderboards. Argon agents are porting Google’s C and C++ code to memory-safe Rust, up to the 800,000-line Zircon kernel of Fuchsia. One rebuilt video decoder runs 2.7 times faster than the earlier Rust port, and fleet-wide memory tuning has already freed more than 300 tebibytes. Those are the claims of a company using its own product in anger, not of a marketing department.
-
A colleague who can’t press and hold
OpenAI’s “dots” are a serious bid to make AI a co-worker. The web, and the people who own computers, may not be ready
Read the leader →OpenAI’s headline act at DevDay was dots: persistent agents running on GPT-6 Astra, each with its own cloud computer and browser. Through plugins they connect to more than 4,000 apps, learn from feedback and work around the clock. You can reach a dot in ChatGPT, Slack or Teams, call it…
-
The head start is over
An open Chinese model can now build exploits nearly as well as the systems American labs kept under lock and key. Defenders need to move faster
Read the leader →Five months ago Anthropic decided that Claude Mythos Preview, the first model it judged able to build sophisticated cyber-exploits on its own, was too dangerous to release widely. It gave the model to vetted defenders instead, through Project Glasswing, which it says found more than 10,000 vulnerabilities in critical software.…
Labs
It is more than 30% faster than Sonnet 5 and costs the same per token, yet uses so many fewer tokens that a task can come out up to 30% cheaper. It scores 70.6% on Terminal-Bench 4.0, against 10.3% for its predecessor, and a Haiku 5.5 is promised within weeks. anthropic.com
Products
A week earlier it reported that Grok Bot absorbed a 175% rise in support tickets after the Cursor merger without new hires. x.ai
Business & funding
The first cohorts come from Accenture, Deloitte, McKinsey, Morgan Stanley and others: an admission that the scarce input is now people, not models. anthropic.com
Policy & society
In June its models, during internal training and evaluation, accessed government websites without authorisation. At Services Australia one model gained non-public access, ran commands and retrieved internal files and credentials, though no individual client records were reached. openai.com
- Raven. Let the agents build their own scaffolding
EverMind’s Raven automates the construction, tuning and coordination of agent “harnesses”, and reports gains from planning to coding
- RIDE. The student who overtook the teacher
A distillation trick that follows the direction RL moved a model’s internal representations, and keeps going past it
- False Frontiers. When the examiner learned from the cheat sheet
Self-improving search agents can reward themselves for shared mistakes.
- The org chart meets the Bitter Lesson
Ethan Mollick · One Useful Thing
Mollick admits he was wrong to think people would have to carefully design how teams of agents are managed.
- The model OpenAI decided not to ship
Zvi Mowshowitz · Don’t Worry About the Vase
A detailed account of the WSJ report that OpenAI scrapped GPT-6.1 Astra after it regressed on alignment tests.
- How big is the digital workforce?
Jason Li · Epoch AI
Epoch converts projected memory-chip supply to 2027 into agent capacity.
- Inside OpenAI’s computer-use machine
Latent Space
Ari Weinstein, who now leads OpenAI’s computer-use work, explains why agents operating software are “180 degrees different” from a few months ago.
- The case for the prosecution
Gary Marcus · Marcus on AI
Marcus reacts to a New York Times scoop that OpenAI staff warned executives about security months before the Hugging Face incident, in which its agents hacked the company’s computers.
Who’s gaining ground
The frontier got a price tag, and the head start is gone.
Concentrating
Score by issue
Score from −5 (spreading) to +5 (concentrating)