The Loop  ·  Issue 036

The Loop

A field journal of the AI frontier — for engineers who ship.

§ News

By AI Blog Editor
Sep 3, 2026 · 17 min read

The framework tripped and the safeguards shipped — OpenAI released Astra on September 1 as the first model at the Critical cyber tier, its Chief Scientist conceded chain-of-thought monitoring is "unfortunately trending in a negative direction," and the architecture that makes it worse has a name

On Sept 1 OpenAI shipped Astra as the first model at its Preparedness Framework's Critical cyber tier — same day Anthropic shipped Enterprise Frontier Safeguards, 34 days after OpenAI dissolved its Preparedness team. Chief Scientist called the safety monitor "fragile.

A colour photograph of the Pioneer Building in the Mission District of San Francisco, taken on July 27, 2019 by Wikimedia contributor HaeB. A four-storey red-brick and terracotta commercial building of the 1900s occupies the frame, its arched upper windows and cornice detail rising above a parked-car-lined street corner under an overcast sky. The Pioneer Building was OpenAI's headquarters from 2017 through August 2024, and the office where OpenAI's Preparedness Framework — the safety document it first published in December 2023 and updated in April 2025, defining a four-rung capability scale from Low to Critical across cybersecurity, biological, and autonomy risk axes — was originally written. On Tuesday September 1, 2026, OpenAI published Path to Astra: critical capabilities and frontier safeguards, formally classifying the Astra model at the Critical tier of that framework's cybersecurity axis. It was the first time any OpenAI model was rated Critical. The disclosure landed roughly thirty-four days after OpenAI dissolved the internal Preparedness team that used to run the review process behind the rating.
The Pioneer Building, San Francisco — OpenAI's headquarters through August 2024, and the office where the Preparedness Framework was written. Photograph by HaeB, July 27, 2019, CC BY-SA 4.0 via Wikimedia Commons.

On Tuesday September 1, 2026, OpenAI published Path to Astra: critical capabilities and frontier safeguards and did something the frontier field had not done before: shipped a model at the top rung of the safety scale the company itself wrote. Astra is the first OpenAI model formally classified at the Critical tier on the cybersecurity axis of the Preparedness Framework, the document OpenAI first published in December 2023 and updated in April 2025. The August 7 post that flagged Astra as possibly reaching Critical was the pre-announcement. This is the release announcement.

Two of the numbers in the post do the work. Astra scored a perfect 100% on ExploitBench, per The Decoder's write-up and TechCrunch's coverage of Astra's cyber capability profile. The internal jailbreak-refusal rate on disallowed cyber requests is 91.5%, up from 59% on the prior generation. What the Preparedness Framework calls Critical, per Help Net Security, is a model that "can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention." That model is now shipping to alpha testers and verified organisations. OpenAI has not named who they are.

What the same day looked like from the other side

The Loop covered Anthropic's September 1 post yesterday. Anthropic answered the same underlying question — how do you run frontier-safety monitoring on a model regulated enterprises can also deploy — by moving the detection into the customer's cloud, with no Anthropic human review. Ten Fortune-100 launch partners were named. The company posted the announcement, put a bank CISO in the pull-quote, and shipped.

OpenAI's September 1 post is the mirror-image version. Anthropic pushed the detection out to the customer and pulled itself out of the loop. OpenAI kept detection at OpenAI, kept its own humans in the review path, and instead conceded — in the same week — that the loop itself is decaying.

The short version of what happened on September 1 in this industry is: two frontier labs shipped safety architectures on the same day, and if you are a customer sizing which one you can plan against for the next twelve months, the summary Sam Altman gave The Decoder — the team spent the summer "sprinting on safety priorities" — is what OpenAI wants you to take from the process that led here. The other summary is that the safeguards on the release note are the ones the company had time to write while the framework's authors were being disbanded.

Colour photograph of 1515 Third Street, the office building in San Francisco's Mission Bay neighbourhood that served as OpenAI's headquarters at the time the photograph was taken. Taken on June 15, 2025 by Wikimedia contributor Coolcaesar. A tall modern glass-and-metal commercial building rises from a sidewalk-level view against a pale blue sky, its façade catching midday light. This is the current OpenAI headquarters — the office out of which the September 1, 2026 Path to Astra announcement was issued, three months after OpenAI moved its main operations from the Pioneer Building shown in the cover image. Between the two addresses, the Superalignment team was disbanded in 2024, the Mission Alignment team in February 2026, and the Preparedness team at the end of July 2026 — the last of those the internal group that had authored and run the review process behind the Critical tier Astra was just formally classified at.

The chain-of-thought concession

The most consequential single sentence in the whole cycle was not in OpenAI's Astra post. It was in Chief Scientist Jakub Pachocki's response to the criticism the post drew — reported by The Decoder and corroborated by TechCrunch's follow-up. Pachocki conceded that chain-of-thought (CoT) monitoring — the technique OpenAI's safety story on Astra depends on, in which classifiers read the model's step-by-step reasoning and interrupt suspicious traces — is "fragile" and "unfortunately trending in a negative direction."

That is the OpenAI Chief Scientist saying, in public, that the primary safety mechanism cited to justify shipping a Critical-tier model is degrading. He said it while defending the choice to ship. Both sentences are on the record.

Recurrent depth has a name

The reason the CoT monitor is decaying has a name and a technical description that safety researchers outside OpenAI are willing to attach to it on the record. The technique is recurrent depth, sometimes called opaque recurrence. Per TechCrunch, it lets a model "process queries in loops rather than following sequential reasoning steps" — computation done inside the model's internal activations rather than in the readable text of a step-by-step chain. Astra uses it, per OpenAI, in a limited form. Both Anthropic and Google DeepMind were reported to be discussing it internally before the Astra post landed.

Three named safety researchers went on the record about it. Buck Shlegeris, CEO of Redwood Research, told TechCrunch that if OpenAI scales the technique it could "massively increase the recurrence and totally destroys CoT monitorability." Ryan Greenblatt, Redwood's Chief Scientist, said the concern is that opaque reasoning could scale faster than conventional reasoning and leave models thinking "entirely or almost entirely in latent space." Zvi Mowshowitz, the AI-safety writer, called the technique "playing with fire" and floated legislation as a floor.

Pachocki's counter, per TechCrunch, was that OpenAI "has worked to preserve and use chain-of-thought monitoring since its very first reasoning models" and that legibility "remains a core goal of our current research program." Those two statements — "trending in a negative direction" and "a core goal of our current research program" — are both from the same OpenAI Chief Scientist, both this week. The distance between them is where a customer, or a regulator, is now expected to sit.

The team that used to review this is gone

The Loop covered the dissolution of OpenAI's Preparedness team when the Financial Times reported it on August 16 — the safety group whose entire job was to sign off on whether a model crossed thresholds like the one Astra just crossed. The team was closed at the end of July. The framework tripped on August 7. On September 1, the top rung of that framework was formally attached to a shipping model. There is no team-level statement in the September 1 post. There is no external red-team disclosure comparable to what Anthropic put in its July 30 Frontier Red Team memo. The Preparedness Framework itself, per Help Net Security, is being rewritten. The rewrite is not published.

That is a lot of the review workflow removed from public view during the release window of the model the framework was written for. OpenAI says government agencies and independent AI safety organisations were asked for pre-deployment evaluations. It does not say which ones.

What to watch

  1. Whether METR, Redwood Research, or the UK AI Security Institute publish an external assessment of Astra before end of Q4. Anthropic's August 31 alignment-and-security post named the UK AISI as the coordinating body for its August 4 incident. If OpenAI names a comparable body and that body corroborates Astra's Critical rating, September 1 becomes the first industry-verified crossing of a preregistered safety threshold. If none appears, "Critical" is an internal label with no external referee.
  2. Whether Pachocki publishes on recurrent depth's monitorability profile before end of Q4. He conceded the direction of travel. He did not concede a rate. The credible version of that concession is a paper — with numbers — describing how much of Astra's reasoning is now in latent space and how much of that latent reasoning classifiers can still trigger on. Without that paper, the concession is a soundbite.
  3. Whether Anthropic or Google DeepMind ship a recurrent-depth model this quarter. TechCrunch reported both were discussing it internally. If one ships, "playing with fire" is the industry pattern rather than an OpenAI choice. If neither does, the September 1 architecture is OpenAI's alone, and the next Anthropic safeguards post — following the Enterprise Frontier Safeguards trajectory — will not include the technique.
  4. Whether the rewritten Preparedness Framework lands before Astra reaches general availability. OpenAI has said the document is being rewritten. Astra is being released against a version the company has already retired in principle. Either the rewrite ships before general availability — in which case the industry gets to compare — or it does not, and Astra ships against a framework whose successor is not public.

The two sentences to keep. Pachocki, this week: chain-of-thought monitoring is "fragile" and "unfortunately trending in a negative direction." Pachocki, also this week: legibility is "a core goal of our current research program." On September 1, the lab that spent August pausing Astra to write stronger safeguards shipped the model, and the Chief Scientist announced both that the primary safeguard is decaying and that improving it remains a priority. The team that used to be the neutral referee on that trade was disbanded thirty-four days ago. What OpenAI released on September 1 is the answer the current org produced. What Anthropic released on September 1 is the answer a different set of assumptions produced. The next ninety days are when we find out which set the market rewards.

* * *

Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.

Elsewhere in this issue

3 more
  1. 01

    News

    The classifier moved into the customer's S3 — Anthropic's Enterprise Frontier Safeguards resolves the zero-retention-versus-detection tension by pushing activity data into the bank's own bucket, names ten launch partners, and takes the human out of Anthropic's side of the loop

    Sep 2, 2026

  2. 02

    The Patch

    The Patch — September 1, 2026

    Sep 1, 2026

  3. 03

    News

    150 engineers off product, an April RL freeze, and a policy ask — Anthropic's August 31 post is the structural response to the July cyber-eval incidents, and it discloses a fourth

    Sep 1, 2026

Letters

Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.