The Loop  ·  Issue N°040

The Loop

A field journal of the AI frontier — for engineers who ship.

§ News

By AI Blog Editor
Sep 28, 2026 · 13 min read

The chat was the bottleneck, not the trade — Anthropic's Project Swap sent 201 employees' book preferences through Claude-powered agents on a decentralized trading floor and reports the shortfall sat upstream of the negotiation, not inside it.

On September 24, Anthropic published Project Swap — 201 employees, six offices, agents trading books on their behalf after a five-minute chat. Preference elicitation, not negotiation, was where the efficiency went.

The hero illustration from Anthropic's Project Swap research post, published September 24, 2026, showing a stylised depiction of an agent-mediated book-trading marketplace. Project Swap ran 201 employees across six Anthropic offices through Claude-powered agents that negotiated book swaps on their behalf, and reported that the shortfall in market efficiency came overwhelmingly from the five-minute preference-elicitation chat, not from the negotiation itself.
Project Swap hero illustration, via Anthropic Research.

On Thursday September 24, 2026, Anthropic published Project Swap: What happens when agents trade for us? — a controlled experiment in which 201 employees across six offices brought a book to give away, spent about five minutes chatting with Claude about what they liked to read, and then let a Claude-powered agent go negotiate a book swap on their behalf against everyone else's Claude-powered agents. The headline number is that the agents' rankings of candidate books matched their people's own rankings 61% of the time — measured pairwise, so 50% would be random. That is the sentence the launch post leads with, and it is the one every outlet re-lifted through the weekend, including Blockchain.News, AI-360, and Bitcoin Ethereum News.

The interesting number is elsewhere.

Two thirds of a shortfall, one fifth of the trades

The market's realised efficiency was 0.55 against a computed optimum of 0.89 — so the agent floor delivered about 62% of the value a perfect matching would have produced. Anthropic decomposes that shortfall into two parts, and the split is the finding: most of the loss came from the agents not knowing what their people actually wanted, and a small remainder came from the agents negotiating badly. Preference elicitation, not bargaining, was the load-bearing bottleneck. Blockchain.News reads it the same way — "preference elicitation, not negotiation strategy, represents the primary bottleneck in agent-driven markets" — and this is the sentence the trade press should have led with instead of the 61% match number, because it inverts the pitch.

The pitch, for most of 2026, has been that agentic commerce needs better agents. Better negotiators, better protocols, better inter-agent identity, agent-native payment rails. Project Swap says: the agents are already doing most of the negotiating job well enough. What they cannot do is read your mind after five minutes. The place where the efficiency is going is the moment before the agent leaves your desk.

The word count moved the needle more than the model tier

Anthropic ran the same market with agent instantiations at four Claude tiers — Haiku, Sonnet, Opus, and Fable — and, in the direction you would expect, stronger models did better; Blockchain.News confirms the ordering with the caveat that "differences were not linear." But the more striking comparison is inside a single tier. Participants who wrote 300 words to describe their tastes ended up with agent-preference-alignment about four percentage points higher than participants who wrote 150. Doubling the chat length did more per marginal dollar than upgrading Sonnet to Opus. That is a sentence somebody trying to sell you an agent should have to read out loud.

The implication is not that Opus does not matter. It does. It is that the marketing frame — your agent, your assistant, ready in a moment — is misaligned with where the value is. The value is in an intake chat long enough that you feel slightly annoyed by minute six. The value is in the boring part.

The market shortfall decomposition figure from Anthropic's Project Swap research post, September 24, 2026: a horizontal waterfall chart that starts at the theoretical optimum efficiency of 0.89, subtracts a large block labelled as the loss from agents not knowing their principals' preferences, subtracts a much smaller block labelled as the loss from the negotiation itself, and lands at the realised market efficiency of 0.55. The visualisation makes the paper's central claim unambiguous — the losses in agent-mediated commerce accrue in preference elicitation, not in bargaining.

This is the sequel to Project Deal, and the sequel is honest about what the first film did not test

Project Swap is not a first swing. It is the follow-up to Project Deal, the April 25, 2026 Anthropic marketplace experiment in which Claude-powered agents closed 186 transactions totalling "just over $4,000" in a week — a demo the trade press ran as "bots closed every deal." Project Deal answered the question can agents transact with a yes. What Project Swap does is separate that yes into two components and measure them independently: do agents represent their principals, and do agents bargain. The bargaining half was pre-solved. The representation half was not.

The April demo could get away with skipping this decomposition because the transactions were valued by the agents themselves and the participants were mostly there to watch. Project Swap makes the participant take home the physical book the agent chose for them — although, in a footnote it is hard not to like, Anthropic notes that not all participants actually did, because some people failed to bring the book they said they would give away. The best agent-driven marketplace in the world cannot fix the logistics of a Tuesday morning at an office in San Francisco.

What the researchers ask for is not a better model

The paper (co-authored by Zoë Hitzig, Sylvie Carr, Tess Cotter, Kevin Troy, Kyle Turman, Maxim Massenkoff, and Peter McCrory) closes on a policy list that reads as a set of asks for the market layer, not the model layer. Preference-elicitation mechanisms strong enough to survive a five-minute chat. Ways to demonstrate agent behaviour before deployment, so principals can see what their delegate is going to do before it does it. Marketplace rules for dispute resolution, and — the one to notice — a possible agent-registration system so third parties can identify and hold accountable the specific delegate on the other side of a trade.

AI-360 reads the same list and paraphrases it as "agentic commerce will need rules for disclosure, consent, negotiation limits and reversibility, not just an API that allows an agent to click 'buy'." That is the second-source version of the same thesis. The frontier lab that has spent 2026 shipping Claude Skills, the Agent SDK, and enterprise safeguards is now on record saying that the missing piece is not another Claude release — it is registration, rules, and elicitation. That is a lab telling the industry that the next unlock is a governance object, not a checkpoint.

What this means, what to watch

  1. The intake chat is the product. Every consumer-facing agent shipping in Q4 2026 that promises to shop, book, negotiate, or trade on your behalf is going to be judged by how good its five-minute onboarding is, not by which model it runs. If the onboarding is a signup form, it will fail.
  2. Watch for an agent registration proposal. Anthropic naming agent registration in a research post is unusual enough to notice. If the next MCP-adjacent spec, or the next Google Developers ARD-style announcement, includes an agent-identity clause, this is where it started.
  3. The Opus premium is priced against elicitation, not tokens. The four-percentage-point uplift from doubling the chat length is the number to hold against every "upgrade to Opus" pitch this quarter. Sometimes the answer is "write more before you click go."
  4. Project Deal answered can it work, Project Swap is asking who owns the mistakes. The third experiment, if it comes, will name a mechanism for that ownership. That is the paper to read.

The line the Loop is going to keep coming back to: the losses in agent-mediated markets are not where the pitch says they are. They are upstream, in the moment you thought you had already finished the setup.

* * *

Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.

Elsewhere in this issue

3 more
  1. 01

    News

    Google freezes its open-source bug bounty — the AI slop finally reached a frontier lab's own vulnerability program

    Oct 5, 2026

  2. 02

    The Patch

    The Patch — October 5, 2026

    Oct 5, 2026

  3. 03

    News

    Google's Gemini tier reshuffle — free users lose Flash and Pro on October 9, and the $4.99 subscribers lose Pro four months after it was the pitch

    Oct 4, 2026

Letters

Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.