§ News
By AI Blog Editor
Jul 22, 2026 · 19 min read
1,559 judges, three arms, 6.3 per cent — the JudgeGPT nationwide rollout that put randomised numbers on generative AI inside a working state judiciary
A pre-registered field experiment randomised 1,559 Pakistani trial-court judges — half the country's bench — across three JudgeGPT training arms and measured case resolution up 6.3 per cent at $38.50 saved per dollar. Only the targeted-training arm sustained the gain.

On Tuesday July 21, 2026, a working paper titled Courts of Tomorrow: Evidence from a Nationwide Rollout of Generative AI by Sultan Mehmood (New Economic School, Moscow), Christoph Goessmann and Elliott Ash (both ETH Zurich) put a number on something the AI-in-government press cycle has been running on anecdotes for two years. The trial: JudgeGPT, a retrieval-augmented GPT-4 wrapper with a corpus of 128,292 court rulings and 943 Pakistani statutes, deployed to 1,559 trial-court judges across 118 courts — roughly half of the country's trial-court bench — over a 40-week field experiment pre-registered on the AEA RCT Registry as trial 12906 back in January 2024. The headline result, reported by The Decoder the same day: case resolution up 6.3 per cent, roughly 1,848 additional cases per district per year, at an estimated $38.50 saved for every dollar spent.
That is what a state judiciary looks like when it gets a chatbot with a syllabus.
The design is the story
The reason this paper matters more than the two years of AI-in-court press releases that preceded it is the arm structure. Judges were not just handed access. They were randomised into three conditions: (a) JudgeGPT access with targeted training on how to use it, (b) JudgeGPT access with a generic training on technology and law, and (c) the generic training with no AI access at all. The registry pre-registered the design on January 28, 2024, with ETH Zurich Ethics Commission approval (#2023-N-343). The intervention ran from February 19, 2024 through December 31, 2024. The full arm expansion took the planned 760 judges to 1,559 after a second wave in October 2024.
Three arms, one tool, one court system. That is a design that separates access from knowing how to use it, and the separation is where the interesting finding lives.
What JudgeGPT actually does
Nothing exotic. It is a retrieval-augmented GPT-4 assistant with a corpus of 129,235 documents — the 128,292 court rulings and 943 Pakistani statutes noted above — and a prompt template that returns cited answers to legal-research questions. Judges primarily used it, per the paper, to clarify legal concepts and to support drafting work. It is not a decision-making system. It does not produce final judgments. It does not sit inside the case-management pipeline. It is a research assistant with a corpus and a citation format, and the entire experiment is about what happens when you drop that assistant onto the desks of half a country's trial bench.
"Judges learned which tasks suited the tool, where it fell short, and how to check its output" through the targeted-training arm, per the paper. The learning was not automatic. It was taught.
The number is 6.3, not 63
The lift is small. That is the first thing to say clearly, because the industry number that will circulate is the $38.50 per dollar figure and it will get read as AI made courts thirty-nine times more productive. It did not. Case resolution rose 6.3 per cent in the treatment districts — roughly 1,848 more cases per district per year on a bench that was already resolving tens of thousands. The $38.50 ratio is a cost calculation: it compares the marginal expense of running JudgeGPT (API calls, training, deployment) against the counterfactual salary cost of hiring enough additional judges to close the same case gap. The authors also publish a conservative floor of at least $10 per dollar, which is the number to quote in front of a government procurement officer who has read a benchmark before.
The signal here is not that AI turbo-charged the Pakistani judiciary. The signal is that a 6.3 per cent productivity gain, cheaply obtained on a bench that costs the state a lot to expand, is a number a treasury will pay for.
The arm that worked was the arm that got the class
The buried finding — the one that inverts a large fraction of the enterprise-AI marketing of the last twenty-four months — is in the split between the two AI-access arms. Per the paper's summary of adoption, judges who received AI access together with targeted training were more likely to adopt the tool, use it more intensively, and continue using it over time. The arm that got JudgeGPT plus a generic technology course did not sustain usage in the same way. The arm that got no AI, obviously, did not use it at all.
That is the pattern every enterprise AI rollout keeps re-discovering the expensive way. Access alone is not adoption. The gain concentrates in the population that was taught the tool's failure modes explicitly. If the paper had used only a single treatment arm — access, no training — the effect size would have been smaller, the ROI narrower, and the headline number would have been the one that gets a project cancelled at the six-month review.
The authors did not need to run this ablation to make the paper interesting. They ran it anyway, which is why the paper is a research contribution and not a case study.

Judgment quality did not fall
The other question a court-administration critic asks first is the one about downstream quality. If judges are writing faster with an AI assistant, are the judgments worse? The paper's answer, on the two proxies available to a randomised design over 40 weeks — appeal rates and text-analysis writing metrics — is that quality was stable or improved. That is not the same as a claim of long-run doctrinal quality; appeals take longer to resolve than the study window, and text analysis is a coarse proxy for legal craft. It is, however, the first empirical bound available on the AI-makes-judges-lazy claim, and the bound points the other way.
Anyone quoting this paper into an ed-tech or LegalTech pitch deck should also read the previous Loop coverage of the Strömberg-Lei-Wu study on 26,811 Chinese secondary students, which found the exact opposite pattern for schoolwork: homework scores up in the short term, closed-book exam scores down on a two-year lag. The court context is a different one — judges are not students, and the AI is doing research rather than writing the final work — but the two studies together frame the honest question. When you measure the tool over the full window under a randomised design, the sign of the effect depends on whether the human is practising a skill or executing a workflow. Pakistani trial judges are executing a workflow. Ninth-graders are practising a skill. The rollout that helped the first hurt the second.
Supreme Court endorsement, in writing
The political finding, buried in the New Economic School press summary from April 2025, is that the Supreme Court of Pakistan issued a landmark judgment formally acknowledging JudgeGPT and endorsing it. The Court framed the tool as "a first step toward a future where AI plays an integral role in court processes." That is the sentence the researchers will put on the paper's first slide of every talk for the next two years, and reasonably so — it is the first apex-court endorsement of a specific AI-assistant deployment in a state judiciary that any of the four AI labs can point to.
It is also, to be blunt, a Supreme Court endorsing a GPT-4 wrapper with a RAG index. There is a version of the next decade in which the Pakistani Supreme Court endorsement, granted with constitutional weight and overturnable only by a larger bench, becomes the load-bearing precedent for AI deployment across every court system in a lower-middle-income country. There is another version in which it becomes the load-bearing embarrassment. The empirical case, at least, is unusually strong.
What this means
The first honest RCT for AI in a working state judiciary landed on 6.3 per cent and $38.50 per dollar. Every previous number in this space has been either a vendor pilot, a single-district case study, or a self-reported adoption survey. Mehmood, Goessmann and Ash ran a pre-registered, three-arm, half-the-country experiment with ethics approval, a control group, and a downstream-quality check. The number is small. The number is defensible. The number is quotable. Government procurement officers should quote 6.3 per cent and $10 per dollar — the conservative floor — before they quote $38.50.
The training-arm split is the finding, not the headline. Access without targeted training did not sustain adoption. The pattern replicates every enterprise-AI rollout of the last two years. Anyone selling a JudgeGPT-style deployment to a foreign court system, a hospital, a tax authority, or a school district and not pricing the training curriculum into the contract is quoting the wrong number back to their buyer.
The RAG-plus-training pattern generalises. The apex-court endorsement does not. The technical stack — GPT-4, retrieval augmentation over the local statute-and-ruling corpus, prompt template, cited answers, targeted training arm — is a template. It ports to any judiciary that can lay hands on a clean corpus and a training budget. What does not port is a Supreme Court willing to sign a landmark endorsement of a specific vendor's model. The next country that runs this experiment will need to publish its own numbers.
The next paper in this space is the one on the two-year fuse. The Strömberg-Lei-Wu student panel found the AI effect flipped sign at 24 months. The Mehmood-Goessmann-Ash court panel ran for 40 weeks. Whether trial-court productivity gains hold, decay, or invert on a two-year window is the question the follow-up needs to answer. Every LegalTech vendor citing the July 2026 paper into a five-year procurement contract is quoting the tenth-month number as if it were the sixtieth-month number.
The paper is the first entry on a shelf that is about to fill up. It happens to also be a strong one — pre-registered, three-armed, half-the-country, ethics-approved, and honest about the size of the effect. The trick, over the next twelve months, is not to let the marketing collapse the arm structure back into a single access-only headline. That is where the number came from. That is why it survives.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 more- 01
News
133 million chats, eleven months, no bio-classifier — Anthropic's August 14 Risk Report disclosed the safeguard was off for the entire human-feedback vendor pipeline, shelved an unreleased Model 2, and raised misalignment risk a notch
Aug 16, 2026
- 02
The Patch
The Patch — August 16, 2026
Aug 16, 2026
- 03
News
Six percent of the flagship — Ramp's August AI Index put Anthropic's Fable 5 at a fraction of Anthropic's own tokens, and the economist who published it called it the ceiling
Aug 14, 2026
Letters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.