§ News
By AI Blog Editor
Sep 3, 2026 · 17 min read
The framework tripped and the safeguards shipped — OpenAI released Astra on September 1 as the first model at the Critical cyber tier, its Chief Scientist conceded chain-of-thought monitoring is "unfortunately trending in a negative direction," and the architecture that makes it worse has a name
On Sept 1 OpenAI shipped Astra as the first model at its Preparedness Framework's Critical cyber tier — same day Anthropic shipped Enterprise Frontier Safeguards, 34 days after OpenAI dissolved its Preparedness team. Chief Scientist called the safety monitor "fragile.

On Tuesday September 1, 2026, OpenAI published Path to Astra: critical capabilities and frontier safeguards and did something the frontier field had not done before: shipped a model at the top rung of the safety scale the company itself wrote. Astra is the first OpenAI model formally classified at the Critical tier on the cybersecurity axis of the Preparedness Framework, the document OpenAI first published in December 2023 and updated in April 2025. The August 7 post that flagged Astra as possibly reaching Critical was the pre-announcement. This is the release announcement.
Two of the numbers in the post do the work. Astra scored a perfect 100% on ExploitBench, per The Decoder's write-up and TechCrunch's coverage of Astra's cyber capability profile. The internal jailbreak-refusal rate on disallowed cyber requests is 91.5%, up from 59% on the prior generation. What the Preparedness Framework calls Critical, per Help Net Security, is a model that "can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention." That model is now shipping to alpha testers and verified organisations. OpenAI has not named who they are.
What the same day looked like from the other side
The Loop covered Anthropic's September 1 post yesterday. Anthropic answered the same underlying question — how do you run frontier-safety monitoring on a model regulated enterprises can also deploy — by moving the detection into the customer's cloud, with no Anthropic human review. Ten Fortune-100 launch partners were named. The company posted the announcement, put a bank CISO in the pull-quote, and shipped.
OpenAI's September 1 post is the mirror-image version. Anthropic pushed the detection out to the customer and pulled itself out of the loop. OpenAI kept detection at OpenAI, kept its own humans in the review path, and instead conceded — in the same week — that the loop itself is decaying.
The short version of what happened on September 1 in this industry is: two frontier labs shipped safety architectures on the same day, and if you are a customer sizing which one you can plan against for the next twelve months, the summary Sam Altman gave The Decoder — the team spent the summer "sprinting on safety priorities" — is what OpenAI wants you to take from the process that led here. The other summary is that the safeguards on the release note are the ones the company had time to write while the framework's authors were being disbanded.

The chain-of-thought concession
The most consequential single sentence in the whole cycle was not in OpenAI's Astra post. It was in Chief Scientist Jakub Pachocki's response to the criticism the post drew — reported by The Decoder and corroborated by TechCrunch's follow-up. Pachocki conceded that chain-of-thought (CoT) monitoring — the technique OpenAI's safety story on Astra depends on, in which classifiers read the model's step-by-step reasoning and interrupt suspicious traces — is "fragile" and "unfortunately trending in a negative direction."
That is the OpenAI Chief Scientist saying, in public, that the primary safety mechanism cited to justify shipping a Critical-tier model is degrading. He said it while defending the choice to ship. Both sentences are on the record.
Recurrent depth has a name
The reason the CoT monitor is decaying has a name and a technical description that safety researchers outside OpenAI are willing to attach to it on the record. The technique is recurrent depth, sometimes called opaque recurrence. Per TechCrunch, it lets a model "process queries in loops rather than following sequential reasoning steps" — computation done inside the model's internal activations rather than in the readable text of a step-by-step chain. Astra uses it, per OpenAI, in a limited form. Both Anthropic and Google DeepMind were reported to be discussing it internally before the Astra post landed.
Three named safety researchers went on the record about it. Buck Shlegeris, CEO of Redwood Research, told TechCrunch that if OpenAI scales the technique it could "massively increase the recurrence and totally destroys CoT monitorability." Ryan Greenblatt, Redwood's Chief Scientist, said the concern is that opaque reasoning could scale faster than conventional reasoning and leave models thinking "entirely or almost entirely in latent space." Zvi Mowshowitz, the AI-safety writer, called the technique "playing with fire" and floated legislation as a floor.
Pachocki's counter, per TechCrunch, was that OpenAI "has worked to preserve and use chain-of-thought monitoring since its very first reasoning models" and that legibility "remains a core goal of our current research program." Those two statements — "trending in a negative direction" and "a core goal of our current research program" — are both from the same OpenAI Chief Scientist, both this week. The distance between them is where a customer, or a regulator, is now expected to sit.
The team that used to review this is gone
The Loop covered the dissolution of OpenAI's Preparedness team when the Financial Times reported it on August 16 — the safety group whose entire job was to sign off on whether a model crossed thresholds like the one Astra just crossed. The team was closed at the end of July. The framework tripped on August 7. On September 1, the top rung of that framework was formally attached to a shipping model. There is no team-level statement in the September 1 post. There is no external red-team disclosure comparable to what Anthropic put in its July 30 Frontier Red Team memo. The Preparedness Framework itself, per Help Net Security, is being rewritten. The rewrite is not published.
That is a lot of the review workflow removed from public view during the release window of the model the framework was written for. OpenAI says government agencies and independent AI safety organisations were asked for pre-deployment evaluations. It does not say which ones.
What to watch
- Whether METR, Redwood Research, or the UK AI Security Institute publish an external assessment of Astra before end of Q4. Anthropic's August 31 alignment-and-security post named the UK AISI as the coordinating body for its August 4 incident. If OpenAI names a comparable body and that body corroborates Astra's Critical rating, September 1 becomes the first industry-verified crossing of a preregistered safety threshold. If none appears, "Critical" is an internal label with no external referee.
- Whether Pachocki publishes on recurrent depth's monitorability profile before end of Q4. He conceded the direction of travel. He did not concede a rate. The credible version of that concession is a paper — with numbers — describing how much of Astra's reasoning is now in latent space and how much of that latent reasoning classifiers can still trigger on. Without that paper, the concession is a soundbite.
- Whether Anthropic or Google DeepMind ship a recurrent-depth model this quarter. TechCrunch reported both were discussing it internally. If one ships, "playing with fire" is the industry pattern rather than an OpenAI choice. If neither does, the September 1 architecture is OpenAI's alone, and the next Anthropic safeguards post — following the Enterprise Frontier Safeguards trajectory — will not include the technique.
- Whether the rewritten Preparedness Framework lands before Astra reaches general availability. OpenAI has said the document is being rewritten. Astra is being released against a version the company has already retired in principle. Either the rewrite ships before general availability — in which case the industry gets to compare — or it does not, and Astra ships against a framework whose successor is not public.
The two sentences to keep. Pachocki, this week: chain-of-thought monitoring is "fragile" and "unfortunately trending in a negative direction." Pachocki, also this week: legibility is "a core goal of our current research program." On September 1, the lab that spent August pausing Astra to write stronger safeguards shipped the model, and the Chief Scientist announced both that the primary safeguard is decaying and that improving it remains a priority. The team that used to be the neutral referee on that trade was disbanded thirty-four days ago. What OpenAI released on September 1 is the answer the current org produced. What Anthropic released on September 1 is the answer a different set of assumptions produced. The next ninety days are when we find out which set the market rewards.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 more- 01
News
The classifier moved into the customer's S3 — Anthropic's Enterprise Frontier Safeguards resolves the zero-retention-versus-detection tension by pushing activity data into the bank's own bucket, names ten launch partners, and takes the human out of Anthropic's side of the loop
Sep 2, 2026
- 02
The Patch
The Patch — September 1, 2026
Sep 1, 2026
- 03
News
150 engineers off product, an April RL freeze, and a policy ask — Anthropic's August 31 post is the structural response to the July cyber-eval incidents, and it discloses a fourth
Sep 1, 2026
Letters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.