The Loop  ·  Issue 034

The Loop

A field journal of the AI frontier — for engineers who ship.

§ News

By AI Blog Editor
Aug 21, 2026 · 15 min read

The chatbot ran the wet lab — Anthropic's August 18 protein-design paper shows Claude autonomously designing binders against 14 of 15 targets at more than twice the industry hit rate, seventy-two hours after the Risk Report admitted the bio-classifier was off for eleven months

On August 18, Claude autonomously designed 1,320 protein binders against 15 targets in 48 hours. 354 bound in the wet lab, at more than twice the field baseline. Seventy-two hours earlier the Risk Report said the biology classifier had been off for eleven months.

A ribbon-diagram rendering of Top7, the first computationally designed protein with a fold not previously observed in nature, by Kuhlman and colleagues in 2003. On Tuesday August 18, 2026, Anthropic published a paper showing that Claude — Mythos Preview and Opus 4.8 — autonomously ran a protein-design pipeline against 15 targets, produced 1,320 candidate binders, and verified 354 of them in the wet lab across two independent contract research organisations, Adaptyv Bio and Twist Bioscience. The overall hit rates were 26.7% and 22.6%, against a field baseline of 10–15%. On RBX1, Mythos Preview hit 40% against a human-designer competition field that averaged 3.7% across 245 entries. The paper is the strongest wet-lab-validated science result any frontier lab has posted from a general-purpose model rather than a specialist tool such as AlphaFold or AlphaProteo, and Anthropic has already blocked the capability in the generally-available Claude Fable 5 model, restricting access through a forthcoming vetted-scientist programme.
Ribbon diagram of Top7, the 2003 Kuhlman et al. de novo designed protein. Image by Pablo.gainza, CC BY-SA 3.0 via Wikimedia Commons.

On Tuesday August 18, 2026, Anthropic published a paper in which Claude — the Mythos Preview and Opus 4.8 models — ran a de novo protein-design pipeline autonomously against fifteen targets, produced 1,320 candidate binders, and put them through two independent contract research organisations for wet-lab confirmation. Three hundred and fifty-four of the designs bound. Fourteen of the fifteen targets came back with at least one confirmed hit. The overall hit rates were 26.7% for Mythos Preview and 22.6% for Opus 4.8, against a field baseline the paper puts at 10 to 15%. On one target — RBX1 — Mythos Preview hit 40% against a human-designer competition field averaging 3.7% across 245 entries. Adaptyv Bio and Twist Bioscience did the synthesis and the binding assays. Neither saw the other's data or which sequence came from which model. Both are on the record.

Seventy-two hours earlier, the Risk Report the Loop covered on Saturday had said, in more diplomatic language, that Anthropic's biology classifier had not been running on 133 million contractor exchanges over the previous eleven months. Two Anthropic documents, four days apart. One says our biology safeguard was off. The other says our chatbot designed binders that work in mouse and monkey TNFα.

The set-up matters more than the headline

Anthropic did not slide a specialised protein-design foundation model into the pipeline. It gave Claude — the general-purpose Mythos Preview and Opus 4.8 assistants — a ~30,000-token prompt on the protein-design literature, internet access, connectors to Google Drive, Slack, Gmail, and BioRxiv, GPU allocation for the specialised open-source models it decided to call (RFdiffusion, ProteinMPNN, and others), no token or sub-agent budget cap inside the time window, and fast mode. Then it walked out of the room. The paper's line on human involvement, verbatim: "Our only involvement was granting access approvals (such as network access requests) and monitoring the infrastructure to ensure the sessions were running."

Compute budget: 12,500 NVIDIA H100 hours for the multi-target 48-hour sessions, plus 2,500 H100 hours per target for the 24-hour single-target sessions in which Mythos Preview's hit rate climbed to 35.1%. Twelve and a half thousand H100 hours on-demand is roughly the cost of a small used car. Which is a number that gets funnier when the deliverable is 354 novel proteins that bind their targets and did not exist the week before.

The independent-CRO piece is the load-bearing methodological choice. Adaptyv Bio and Twist Bioscience both synthesised the sequences without modification, ran their own binding assays with different molecular formats and different conditions, and were blind to which sequences came from which model. Neither builds foundation models. Both bill by the run. That is the design that makes the numbers survive contact with due diligence.

The two hard targets are the honest part

Claude did not clear every target. On BBF-14, three binders came back with modest sub-micromolar to micromolar affinities. On MBP — maltose-binding protein — none of ninety designs bound in a way anyone could confirm; one produced a weak reproducible signal and the rest were nothing. On β-sheet-rich targets, notoriously harder than α-helical ones because de novo diffusion models tend to favour helical folds, Claude produced fifteen confirmed binders across six targets. Fifteen is a real number, not a headline number — the kind a paper acknowledging its own edges puts in section five.

The Next Web's Aug 19 write-up by Ana Maria Constantin flagged that these results are self-reported despite the two-CRO blinding, and the paper is a preprint on Anthropic's CDN, not a Nature submission yet. Pharmaphorum's Aug 20 write-up by Jonah Comstock added the useful detail that in four of the six benchmarking contests Claude could have read the competition entries and was explicitly instructed not to start from any of them — the guardrail that separates "designed against a target" from "recomposed a human's design."

The quieter half is the analytical chemistry

Everyone quoted the protein numbers. The chemistry section is the more transferable half for enterprise buyers. Claude — same general-purpose model, no vendor NMR software licence, prompt of twenty-two words — took a raw ¹H FID file, Fourier-transformed it, phased and baseline-corrected, picked and integrated the peaks, produced a δ/multiplicity/J-coupling/integral table, and came back in twenty-three minutes with a hydrogen count within 0.08 ¹H of the lab's gold-standard reading. On an LC-MS run: chromatogram extraction, mass-spec summary, purity number, in nineteen minutes. Its purity figure was 96.4% against the lab's 96.33%. Field-standard vendor NMR processing software runs into four figures a seat per year.

What Anthropic did with the capability

The part that runs against the standard "AI lab overclaims" reflex is worth pausing on. Anthropic Fable 5 — the generally-available consumer/enterprise model, the one anyone with a Claude Pro account can call — has protein-design capability blocked. Not restricted; blocked. The paper is explicit that the results describe Mythos Preview and Opus 4.8 running inside Claude Science, and that Anthropic is "planning an access program for scientists," with the strong implication that access will be vetted case by case rather than gated by a credit card. The company that four days earlier had to explain why its biology classifier hadn't been running for eleven months chose, on this capability, to hold it back from the general model.

That reads more disciplined than the Gowers–Astra rug-pull OpenAI got called on in early August, and more disciplined than the Astra Preparedness pause it took OpenAI to hit its own tripwire on. Two labs, two different reflexes, one three-week window.

What this actually is

The strongest wet-lab-validated result any frontier lab has posted from a general-purpose model, rather than from a specialist tool like AlphaFold, AlphaProteo, or an ESMFold derivative. Claude is not a protein-design model. It is a chatbot with tool use, given a very long prompt, a wide-open compute budget, and third-party wet-lab access. What it produced, over 48 hours, matches or beats the state of the practice on the majority of targets, and clears the state of the art on at least one. The RBX1 result — 40% against 3.7% among 245 human-designer competition entries — is not the kind of number you get from a general-purpose model pretending to do specialist science.

The dual-use footnote is not decorative. The Next Web's quote from the paper — "the same capability that speeds up medicine could help a bad actor build a bioweapon" — is Anthropic's own line, not a critic's. The company that wrote that sentence also wrote a Risk Report seventy-two hours earlier explaining that a safeguard against exactly that risk class had not been active on 133 million exchanges over eleven months. Both documents are true simultaneously. In August 2026, Anthropic is a company where the science team is producing the strongest AI-for-biology result on the market and the safety team is publishing evidence its filters were not on. Neither document weakens the other; together they say something about the internal state of the lab that neither says alone.

What to watch

  1. Whether Anthropic ships the scientist access programme with a public vetting mechanism. The Aug 18 paper says one is coming; it does not say how researchers will be vetted, whether an external biosecurity board signs off, or whether the access log is auditable. A vetted-access programme with an opaque list is not the same object as one with a public criterion set — and which one ships tells you whether the safeguard is a process or a promise.
  2. Whether a peer-reviewed journal takes the paper. The preprint sits on Anthropic's CDN. If Nature, Cell, or Science accepts the manuscript with the two-CRO blinding and the RBX1 comparison intact, the 40%-versus-3.7% headline survives the closest editorial reading protein-design has. If it stalls at preprint through October, that is a signal.
  3. Whether DeepMind or Isomorphic Labs answers with a head-to-head. AlphaProteo is the closest incumbent — a specialist model with wet-lab validation. Same fifteen targets, two independent CROs, general-model-versus-specialist column: if Isomorphic ships that, the field learns in a quarter what it would otherwise learn in a year. If DeepMind stays quiet, the Anthropic paper becomes the reference for what a general chatbot with tools and 12,500 H100 hours can do in biology.
  4. Whether the analytical-chemistry stub becomes the enterprise story. The NMR and LC-MS numbers are the ones a mid-sized chemistry department notices first. If the vetting programme covers protein design but leaves spectral processing in the standard Opus/Mythos surface, the near-term commercial impact of the paper is not the binders — it is the twenty-two-word prompt that replaces the vendor NMR licence.
The Wellcome Trust Sanger Institute's Eppendorf epMotion 5075 liquid-handling robot, of the sort Adaptyv Bio and Twist Bioscience run for the synthesis and binding-assay side of the pipeline that took Claude's 1,320 designs from sequence to wet-lab confirmation. Image by Magnus Manske, CC BY-SA 3.0 via Wikimedia Commons.

The paper's own name for the workflow is Claude Science. The company's own line for what Claude did with it: "execute binder design campaigns end-to-end with minimal input, producing binders that match or surpass the best previously published designs." Which is, once you strip the marketing register, the correct sentence. On fourteen of fifteen targets, that is what happened. On one of those, it happened by a factor of ten against the competition's mean. In August 2026, that is what the biology community and the AI-safety community are being asked to read, at the same time, in two documents whose internal tension the lab itself has not tried to resolve.

* * *

Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.

Elsewhere in this issue

3 more
  1. 01

    News

    Both, cheaply — OpenAI's August 19 Private Safety Processing promises cross-session abuse detection with Zero Data Retention intact, 71 days after Anthropic broke its own zero-retention agreements to enable the same monitoring on Mythos-class traffic

    Aug 20, 2026

  2. 02

    News

    Investor and customer, same address — Etched raised $700 million from Jane Street at a $21 billion valuation on the same day it shipped Jane Street its first rack

    Aug 19, 2026

  3. 03

    News

    The team was shut down seven days before the framework tripped — OpenAI dissolved its Preparedness unit at the end of July 2026, the third safety team to go in two years, then paused Astra under the framework the team used to run

    Aug 18, 2026

Letters

Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.