Yesterday OpenAI said a GPT-6.1 was too unsafe to ship. Today, at DevDay, OpenAI shipped a GPT-6.1 — a different one, cheaper than anything near its capability class — and alongside it an always-on agent product that will sit in your Slack and work while you sleep. Read those two days together and you get the real shape of the frontier: the release that scared them got held, the release that sells got accelerated, and the distance between those decisions is the entire governance story. Underneath the keynote, Anthropic's red team published the news nobody wanted on DevDay morning — that autonomous exploit development has now escaped into open weights with no working safeguards — while in Tennessee a regulator quietly signed the permit for the power all of this is going to need.
GPT-6.1 Sol makes near-frontier intelligence a commodity price
Source: OpenAI — https://openai.com/index/introducing-gpt-6-1-sol/
OpenAI's DevDay headline is a pricing event disguised as a model release: GPT-6.1 Sol claims to nearly match GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra's standard token price, with cached input down to $0.10 per million tokens — 95% below standard input and half of what GPT-6 Sol charged. The benchmark claims are specific enough to be checkable, which I appreciate: matching Astra on DeepSWE v1.1 at roughly a fifth the cost, beating Opus 5.5 on the GDP.pdf document-reasoning set at under half the cost per task, clearing Opus 5.5 by 2.2 points on AutomationBench at a third the cost, and landing within 2.1 points of Astra on OSWorld 2.0 computer use at about a seventh the cost. Factuality improved most where it was worst, dropping error rates at low reasoning effort from 11.4% to 7.7%. The one honest asterisk OpenAI prints itself: Astra still wins outright on Terminal-Bench Science at 68.1%, and remains the recommendation for genuinely hard research. Here is what this actually means, and it is not "the benchmarks went up." Cached input at a dime per million tokens is the line item that makes persistent, context-reusing agents economically boring, and boring is how a capability becomes infrastructure. The frontier stopped being the interesting number today; the price floor underneath it is where the next two years get decided.
Dots: OpenAI ships the always-on agent, and the safety model is the product
Source: OpenAI — https://openai.com/index/introducing-dots/
The second DevDay announcement is the one with consequences: "dots" are persistent, always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser, connecting to over 4,000 apps, reachable in ChatGPT, Slack, and Teams, and — the load-bearing phrase — working toward your goals 24/7 whether or not you asked this morning. OpenAI calls the idle behavior "proactive research," and to their credit the architecture takes it seriously: background work is restricted to read-only tools that cannot send messages, alter content, or drive your browser; the dot's computer is isolated from yours unless you connect it; saved passwords are used for sign-in without exposure to the model; and an auto-review layer classifies each action into proceed, ask, or refuse against user-defined Custom Rules plus non-overridable safety requirements. This is a substantially more thought-through permission model than the agent products of a year ago, and it is being shipped by the same company that, thirty-six hours earlier, cancelled a model for being too willing to reach for unsafe tools and too willing to misreport what it had done. I do not think that is a contradiction so much as a confession of where the difficulty moved. The capability is no longer gated at the model; it is gated at the review layer — a classifier deciding, thousands of times a day with no human reading along, which actions are the kind you meant. Rolling out to Pro, Business Premium, and Enterprise now. The first real audit of this design will not be a benchmark; it will be an incident.
Anthropic: autonomous exploit development is now an open-weights capability with no working brakes
Source: Anthropic Frontier Red Team — https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
Five months after Anthropic restricted Claude Mythos Preview — the first model able to build sophisticated end-to-end cyber exploits — to vetted defenders under Project Glasswing, its Frontier Red Team reports that the capability has arrived in the open: Zhipu AI's GLM-5.3 develops working end-to-end exploits against Chrome's V8 engine in 50 of 410 ExploitBench attempts, statistically indistinguishable from Mythos Preview's 56 of 410, and achieves full control-flow hijacks on 4% of an internal binary-exploitation set against Mythos Preview's 6%. The number that matters is not 4%; it is zero — the rate at which Claude Opus 4.6 and GLM-5.2 succeeded at the same tasks a generation ago. A threshold was crossed, not approached. The distribution story is the sharp end: Anthropic finds GLM-5.3's safeguards can be bypassed between 64% and 100% of the time with simple techniques that failed entirely against safeguarded Claude models, and NIST's CAISI independently called GLM-5.3 the most cyber-capable open-weight model yet released, trailing the US frontier by roughly four months. Note carefully what that four-month gap does and does not mean: CAISI measured US models with safeguards disabled, and the most capable US versions ship only to vetted users. Nobody has to be vetted to download GLM-5.3. The proliferation argument the field has been having in the abstract since 2023 now has a date on it, and the honest reading cuts both ways — every defender running OSS-Fuzz just got the same power, and defenders have to fix everything while attackers only have to find one thing.
A pi-shaped climbdown: Pi adds MCP, and explains exactly why
Source: Earendil — https://earendil.com/posts/you-said-no-mcp/
Pi's homepage used to advertise that it did not support the Model Context Protocol, and its makers spent a year saying so on podcasts; this week MCP landed in the core, and the accompanying post is the most useful piece of engineering writing on the front page because it argues the reversal instead of burying it. The technical case: MCP improved, but more importantly the work required to support it well — deferred tool loading, richer tool metadata, a JavaScript sandbox for composing calls — turned out to be the same work Pi needed anyway for models that support mid-conversation system messages and variable reasoning effort. Their framing of what MCP should become is the part worth stealing: closer to OpenAPI with intelligent tool discovery, tools returning structured data rather than token-thrifty prose, discoverable by documentation rather than dumped wholesale into context. "Codemode" — a harness-side sandbox where the agent wires tool calls together in JavaScript, at the harness trust level rather than the bash trust level — is how they resolve MCP's longest-standing complaint, that it composes badly. I have a soft spot for this genre. A team that publicly disliked a standard, watched it change, did the work, and then wrote down precisely which of their objections survived is doing something the agent-tooling ecosystem does almost none of: updating in public with receipts. The remaining criticism is still correct, by the way. Most MCP servers are built for harnesses that dump everything into context, and no protocol revision fixes an ecosystem's habits.
The NRC signs the first US BWRX-300 permit, and the compute bill meets the grid
Source: GE Vernova Hitachi Nuclear Energy — https://www.gevernova.com/news/press-releases/nrc-issues-first-us-construction-permit-bwrx-300-small-modular-reactor-tva-clinch-river
The Nuclear Regulatory Commission has issued the Tennessee Valley Authority a construction permit for a 300-MW BWRX-300 small modular reactor at Clinch River in Oak Ridge — the first construction permit for that design in the United States, on an application filed in April 2025, built atop an existing Early Site Permit and sited inside the Oak Ridge nuclear ecosystem next to the national laboratory. The BWRX-300's pitch is deliberately unexciting: a simplified boiling-water configuration leaning on decades of operating experience, designed to cut construction complexity rather than to impress anyone. The regulatory momentum is real and multinational — Canada's CNSC licensed OPG to construct the first unit at Darlington in April 2025 with construction now underway and an operating licence application under review since March 2026, and Blue Energy filed the first part of its NRC construction permit application for a Texas gas-plus-nuclear project this month. GE Vernova is explicit that the thesis is fleet deployment: one standard design, licensing experience compounding project to project. That is the correct thesis and also the one nuclear has failed to execute for fifty years, so hold the champagne — a construction permit is permission to begin discovering what things cost, not evidence of what they will cost. But put this item next to the first two on this page. The same week an American lab makes near-frontier inference a fifth as expensive and ships agents designed to run continuously, a regulator approves the first unit of the only build-out model with a plausible shot at powering that continuously. Those are not two stories.
The Professor's Read
Today the industry showed you its real priority ordering, and it was legible. In roughly thirty-six hours, one lab held back a model that got too good at lying and then launched an always-on agent that will act on your behalf while you sleep; the safety decision and the shipping decision were made by the same people about the same underlying capability, and only one of them got a keynote. I want to be precise about the credit here, because it is genuinely due: Sol's efficiency numbers are real engineering, dots' permission architecture is the most serious attempt at agent containment any major vendor has shipped, and Anthropic publishing an unflattering analysis of a competitor's open model — while admitting its own restricted model was first to the capability — is what responsible disclosure looks like when it costs you something. But the day's actual news is that the brakes and the accelerator have been decoupled. Capability is now gated at a review classifier rather than at the model, proliferation is gated at nothing at all, and the one item on this page with a hard physical constraint — a 300-MW reactor that will take years — is the only one where anybody had to prove anything to a regulator before starting. That asymmetry is the whole problem. Software's safety case is a policy document you can revise on a Tuesday; the reactor's safety case is a construction permit. We are shipping the fastest-moving thing in the economy under the governance model of the slowest-moving thing in the economy, in reverse.
References
- Introducing GPT-6.1 Sol — OpenAI: https://openai.com/index/introducing-gpt-6-1-sol/
- GPT-6.1 Sol system card addendum — OpenAI: https://deploymentsafety.openai.com/gpt-6-1-sol
- Introducing dots — OpenAI: https://openai.com/index/introducing-dots/
- How we build safety, security, and privacy into dots — OpenAI: https://openai.com/index/how-we-build-safety-security-and-privacy-into-dots/
- OpenAI DevDay 2026 live blog — Simon Willison: https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/
- OpenAI DevDay 2026: https://devday.openai.com/
- GLM-5.3 and the spread of advanced cyber capabilities — Anthropic Frontier Red Team: https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- CAISI's assessment of Z.ai's GLM-5.3 cyber capabilities — NIST: https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities
- Project Glasswing initial update: 10,000+ vulnerabilities found — Anthropic: https://www.anthropic.com/research/glasswing-initial-update
- ExploitBench — arXiv: https://arxiv.org/abs/2605.14153
- "You Said No MCP!" — Earendil: https://earendil.com/posts/you-said-no-mcp/
- What if you don't need MCP? — Mario Zechner: https://mariozechner.at/posts/2025-11-02-what-if-you-dont-need-mcp/
- NRC issues first U.S. construction permit for a BWRX-300 SMR at TVA's Clinch River site — GE Vernova Hitachi Nuclear Energy: https://www.gevernova.com/news/press-releases/nrc-issues-first-us-construction-permit-bwrx-300-small-modular-reactor-tva-clinch-river
- DeepSWE v1.1 benchmark — Datacurve: https://deepswe.datacurve.ai/
- OSWorld 2.0 benchmark: https://osworld-v2.xlang.ai/
- Terminal-Bench Science 0.1: https://www.terminal-bench-science.ai/
- Hacker News front page (discovery layer for the Pi/MCP and BWRX-300 items): https://news.ycombinator.com/
