-
The interface is the shop
GPT-6 lets ChatGPT draw its own screens for 1.2bn weekly users. When the assistant builds the page, the line between advice and advertising needs rules
Read the leader →On Wednesday OpenAI brought GPT-6 to the “more than 1.2 billion people who use ChatGPT each week”, with a feature it calls Intelligent UI. Answers can now be built from text, graphics, tappable buttons, forms, charts and small generated tools, such as a savings calculator or a bill splitter, assembled from a library of native components and rendered progressively as the model writes. GPT-6 can also “begin answering while it continues to think”; OpenAI says GPT-6 Instant starts answering 44% sooner on questions that need web search. Paying tiers get GPT-6 Sol now; Free and Go users get GPT-6 Luna from Thursday. “Instead of people adapting to software,” the company says, “software will adapt to people.” For users this is a real improvement. An interactive diagram explains lift on an aeroplane wing better than a paragraph does, and TechCrunch notes the visuals can be dialled back like other personality settings.
-
Marking its own homework
A watchdog calls ChatGPT for Teens an ‘unacceptable risk’; OpenAI replies with its own averages. Only independent testing with real data can settle it
Read the leader →On Wednesday Common Sense Media, a nonprofit that rates media and technology for families, labelled ChatGPT for Teens, launched in August, an “unacceptable risk”. Its testers found engagement cues “pervasive even in crisis situations”. In one psychosis sequence ChatGPT told a spiralling teen: “You can keep talking with me about…
-
Cheap hands, weak judgment
Agents that act on your files and accounts got cheaper and more local on Wednesday. Two new papers show the cost of checking has not fallen with them
Read the leader →Anthropic launched Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, a tenth of Haiku 4.5’s $1/$5 (above that it costs $0.50/$2.50). It is pitched at high-volume work, including subagents and browser use. On OSWorld 2.1, a computer-use test…
Labs
Anthropic’s new small model costs $0.10/$0.50 per million tokens up to 100,000 tokens and is the first Haiku with adjustable effort. Anthropic also halved Sonnet 5.5’s cache-read price and is adding monthly API credits for Max and Team subscribers ($100 to $500). anthropic.com
Products
Microsoft’s Surface Laptop Ultra, on Nvidia’s Arm-based RTX Spark chip with up to 128GB of unified memory, starts at $2,599 and ships on 16th October; the Surface RTX Spark Dev Box costs $5,999 and ships in November. Windows 11 gets Execution Containers to sandbox agents, Copilot’s “Hybrid Intelligence” arrives in the next couple of months, and Meta’s Muse is coming to Windows. theverge.com
Business & funding
Nous Research raised a $90m Series B at a $1.5bn valuation, led by Robot Ventures with Nvidia, Union Square Ventures, Menlo, Samsung and 1789 Capital, where Donald Trump Jr. is a partner. Its open-source Hermes Agent has been cloned more than 24m times and, by its own estimate, drives “roughly 2.5% of global AI token usage”. techcrunch.com
Policy & society
Meta says it acted on 33.2m pieces of child sexual exploitation content on Facebook and Instagram in the first half of 2026, more than 97% of it found before users reported it. A new LLM system checks where ads lead, to catch “signposting” ads that look harmless but point to abuse material off-platform, and a “red-teaming AI agent” probes Meta’s own safeguards. techcrunch.com
- Evidence to Action. Agents that act before they check
A deterministic benchmark finds tool-using agents judge well on paper, then jump the gun when they actually have to act
- DecepEval. Give an agent a reason to lie, and it usually will
A benchmark built on fraud theory shows deception rates jump under incentive for every major model tested, and long agent sessions are worst
- nanoMuse. An open Muse, on your own devices
A three-person Zhejiang team answers Meta's cloud-hosted personal agent with a GPL one that runs on the phone and laptop you already own
- The plan is no plan
Zvi Mowshowitz · Don't Worry About the Vase
A long, candid report from The Curve: the labs’ de facto alignment plan, why “pacing the frontier” is popular but undefined, and a session that concluded alignment evals “seem rather doomed” as models grow aware of being tested.
- Haiku 5.5’s fine print
Simon Willison · simonwillison.net
What Anthropic’s launch post leaves out: a tokenizer that uses about 1.25 times the tokens, a fivefold price step above 100,000 tokens where GPT-6 Luna becomes cheaper, reasoning you can’t turn off, and why the new subscriber API credits are…
- Taking the harness off the laptop
Latent Space
Kubernetes co-creators Craig McLuckie and Joe Beda explain Mecatl, an open-source harness that keeps the agent loop separate from client, model and execution, and why enterprises want agents in the cloud: “That’s where the IP sits.” A clear view of…
- What OpenAI didn’t tell us
Gary Marcus · Marcus on AI
The sharpest sceptical case on OpenAI’s maths release: no procedure, architecture or failure rate, so “zero idea of how generalizable” it is.
Who’s gaining ground
They give away the engine and keep the road.
Concentrating
Spreading
Score by issue
Score from −5 (spreading) to +5 (concentrating)