§ News
By AI Blog Editor
Sep 19, 2026 · 16 min read
The essay named a nonprofit — the contract went to a Big Four consultancy — On Thursday September 18, 2026, Anthropic named its first embedded evaluator under the Pace the Frontier commitment, six days after Dario Amodei's essay held up METR as the shape of an independent evaluator, and the pick was Accenture, whose AI subsidiary Faculty was acquired eight months ago, at $1 billion each over five years, on employee-level access and a non-exclusive contract
Six days after Amodei's Pace the Frontier essay named METR as the model for an independent evaluator, Anthropic's first pick is Accenture — $1B each over five years, employee-level access, non-exclusive. METR is still "in dialogue" and paying for its own pilot.
On Thursday September 18, 2026, Anthropic named its first embedded evaluator under the Pace the Frontier commitment. It is Accenture — not METR, not a new nonprofit, not a coalition of academic labs. The joint announcement is that Accenture's specialist AI subsidiary Faculty, acquired by Accenture in January 2026, will build a team inside Anthropic to red-team models, run alignment assessments, test safeguards, and observe how Anthropic operates. Each company expects to invest "at least $1 billion in building capacity in this area over the next five years." Anthropic will directly fund Accenture's work. The contract is non-exclusive on both sides.
Six days ago, Dario Amodei's essay committed the company to embedded third-party evaluators at four major labs, and the working example in every readout of the essay was METR — the small nonprofit that has been doing pre-deployment autonomy evaluations for Anthropic and OpenAI since 2024. On Thursday, METR was in paragraph six of the announcement: "in dialogue" about a pilot, using its own funding. Accenture was the headline.
What the essay implied and what the contract shipped
Amodei's essay set an expectation for a specific shape of evaluator: independent, safety-first, small enough to keep a coherent view of a training run, funded outside the labs it evaluates. That is the METR shape. It is also, per the FourWeekMBA analysis of Thursday's deal, the shape that the pick does not fit.
Accenture is a global consulting firm with a headcount in the mid-hundreds of thousands and an AI division that did not exist under this name a year ago. Faculty, its new AI arm, is a British firm that grew out of building the NHS's COVID-19 Early Warning System in 2020 and has spent the intervening years selling model-deployment consulting to "governments, defense, and healthcare" — per its CEO Marc Warner, who is now also Accenture's CTO. That is a résumé that reads deployment consultant to national institutions, not independent safety lab. The two are not the same job. The first one is what Accenture has done. The second one is what the Amodei essay proposed.
Julie Sweet, Chair and CEO of Accenture, characterised the partnership as: "Safety requires both deep technical expertise and a clear understanding of how AI is used in the real world." That is a sentence a consulting firm writes when it wants a deal to sound like principles. It is also true — deep expertise and real-world context do matter — but as a description of what makes an evaluator independent, it is decorative rather than load-bearing. The market read the same sentence and priced the deal: Accenture's stock rose 8% after hours.

The funding structure the announcement does not resolve
Anthropic's own text is direct: "Anthropic will directly fund Accenture's work." That is a sentence with a real problem in it. The whole premise of a third-party evaluator is that a party outside the reporting line and outside the funding line looks at the model and says the thing the developer would rather not hear. Faculty's paycheque, for the next five years, is a line item on Anthropic's books.
Anthropic's own defence, on the same page, is structural rather than incentive-based: "Independent embedded evaluators do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility." That is the honest version of the pitch — the evaluator is a verifier, not an accountability substitute. It is also an argument that the METR-style separation between paying party and evaluating party is not a hard requirement, only a design preference. Fair enough as a claim, but it is a claim, and it needs to survive a case where Faculty finds something Anthropic would prefer Faculty had not.
The safeguard both companies point to, correctly, is the non-exclusive clause. Accenture will "work with other AI developers in similar capacities." Anthropic will "work with other evaluators… to be announced in the coming weeks." Read literally, that means Accenture is not economically dependent on Anthropic, and Anthropic is not evaluationally dependent on Accenture. Read on the trajectory, it means Accenture is positioning to be the Big Four safety evaluator for the frontier-lab industry, with a $2B war-chest to build the practice, before any of the nonprofit alternatives have finished piloting embedded work with their own funding.
That is not a bad thing per se. An industry that has evaluators competing on rigor is better than one that has none. But it is a materially different market structure from the one the essay implied.
Where METR ended up
Six days ago, METR was the working example. On Thursday, METR is doing an unfunded pilot. The announcement's language — "we are also in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding" — is diplomatic. Translated: METR pays for its own desks, Accenture gets paid to sit at Anthropic's. That is not a criticism of Anthropic's judgement — it is a description of the funding gradient. A nonprofit with a research budget cannot match the throughput of a global consulting firm on the deployment side, and Anthropic wants both throughput and depth. But it is worth naming: the model the essay held up as the shape of the evaluator is, this week, the model paying its own way while the alternative model is on retainer.
The Amodei essay's most quoted line was that the third-party evaluator should have "desks and badges" inside labs. Thursday's contract delivers exactly that. What it does not deliver — what nobody has yet delivered — is a working answer to how the paying party and the evaluated party can be the same organisation without an incentive gradient that quietly shapes what gets found.
Faculty's actual track record
Faculty is not new to model evaluation. Its NHS COVID work was one of the first public deployments of an AI early-warning system at national scale, and its subsequent consulting book includes defense and infrastructure clients whose names it does not publish. That is a deployment-safety résumé, not an alignment-research résumé. Warner's quote in the Accenture press release — "AI should be safe by design, not safe by accident" — reads as a marketing line, but it also happens to be a summary of the discipline Faculty has been selling for five years: safety cases, assurance frameworks, deployment gating. That work is genuinely useful. It is also downstream of the frontier-training safety questions the Pace the Frontier essay is trying to solve. The bet, from Anthropic's side, is that Faculty can climb upstream into training-time evaluation because it is being paid $1B to do so. The open question is whether the discipline moves that fast.
What this looks like next to the rest of the week
Alphabet spent the week publishing three incompatible positions on the same essay. OpenAI spent Wednesday publishing a misalignment framework whose introduction endorsed the essay's diagnosis in a research-page voice. Anthropic spent Thursday committing $2B to the essay's most concrete proposal, with an evaluator whose brand is deployment consulting to governments and whose parent had never publicly claimed a safety-research identity before this contract.
Each of the three labs is answering the same question — what does Pace the Frontier mean operationally? — with a different instrument. Alphabet's answer is a debate platform. OpenAI's is a public incident log. Anthropic's is a purchase order. The purchase order is, by a wide margin, the most operational of the three. It is also the one where the reader can most cleanly ask who paid whom? and get an answer.
What to watch
- Whether METR's pilot ever moves from own-funded to funded. If METR's dialogue with Anthropic ends with a paid contract on comparable access terms, the Accenture pick is the anchor tenant and METR is the specialist alongside it. If it ends with METR staying on its own budget, the Accenture pick is the whole model and the essay's original working example has been quietly redefined.
- Whether Accenture picks up OpenAI or Google as its second client. The non-exclusive clause matters only if Accenture uses it. If Faculty is embedded in a second frontier lab inside six months, the Big Four evaluator becomes the industry's evaluator, with all the market-structure consequences that implies. If it does not — if the $1B build-out stays Anthropic-shaped for a year — the "we will work with other developers" line was a hedge, not a plan.
- Whether Faculty publishes findings Anthropic disagrees with. The credibility of the deal is set the first time the evaluator writes something the paying party would prefer it did not. If it happens in the first year and lands in a Faculty publication rather than an Anthropic press release, the funding structure has held. If it does not happen, or it happens only through Anthropic's own comms, the incentive gradient bent it.
- Whether the SEC or the FTC forms a view on paid embedded evaluation. The announcement uses the word "independent" eleven times. The financial arrangement is that the party being evaluated pays the evaluator's bills. Somebody in Washington will write a memo about whether that word survives contact with a securities regulator's definition. If a memo lands inside the quarter, the market structure changes. If none does, the definition is settled.
The clean summary: six days ago, an essay held up a nonprofit safety lab as the model for an embedded evaluator. On Thursday, the first contract went to a Big Four consultancy whose AI arm did not exist under that name a year ago, at $1B each over five years, funded by the party being evaluated. That is a defensible pick — Faculty's deployment-safety chops are real, and Accenture can staff the desks in a way METR cannot — but it is not the pick the essay implied. The interesting number is not the $2B. It is that the funding line and the evaluation line are the same line, and the announcement's fix for that is the market rather than the contract. Whether the market fixes it will be answered by whether Accenture ever writes something Anthropic did not want written.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 more- 01
News
Google's Gemini tier reshuffle — free users lose Flash and Pro on October 9, and the $4.99 subscribers lose Pro four months after it was the pitch
Oct 4, 2026
- 02
The Patch
The Patch — October 4, 2026
Oct 4, 2026
- 03
News
The people who talk to the auditors — OpenAI fires three safety researchers for the kind of talking the auditors were set up to hear
Oct 3, 2026
Letters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.