Anthropic's economics research team published the results of an internal experiment called Project Swap on September 24. Two hundred and one employees each brought in one physical book. Each first spoke with Claude for about five minutes about what they liked to read, then Claude-powered agents haggled over trades on their behalf in a digital marketplace, and the books that came out of those trades were handed back to their new owners.
Participants came from six offices — San Francisco, New York, London, Seattle, Washington D.C., and Dublin — with 115 from San Francisco and 57 from New York. The market ran for 205 rounds in total, counting the initial run and repeat trials.
How the Market Was Built
The marketplace ran on a decentralized “free trade” model: agents could swap one-for-one or pull several parties into a rotating cycle. Every completed trade was visible to participants and to other agents. The research team randomly assigned agents one of two instructions — either to fight purely for their own owner's interest, or to also weigh other participants' welfare while negotiating.
Two benchmarks were used for comparison: a utility-maximizing optimal allocation, and the Top Trading Cycles rule commonly used in market design.
Stuck at “Reading People”
After the five-minute conversation, Claude ranked all the books for each participant. Compared pairwise against each person's own ranking, the match rate was 61%. For reference, random guessing scores 50%, ranking by popularity gets about 53%, collaborative filtering about 55%, and having a participant's friend guess comes in around 57%.
The trouble showed up downstream. The books participants actually ended up with scored 0.55 on average by their own reckoning — roughly a mid-table pick out of ten books — versus 0.89 for the theoretical optimum. Breaking that gap down, the research team found about 85% came from Claude misreading preferences, and only about 15% came from inefficiency in how agents negotiated.
the market fell short mostly because of the information agents lacked about their participants, rather than because of how they traded
Another data point backs this up: doubling how much text a participant typed raised ranking accuracy by roughly 4 percentage points. The more people talked, the better the agent understood them.
Model Size Matters More Than Instructions
Measuring trading efficiency against Claude's own rankings, Haiku 4.5 scored 0.75, Sonnet 4.5 scored 0.80, Opus 4.8 scored 0.88, and Fable 5 scored 0.86, against a theoretical optimum of 0.95 on this measure. Moving from Haiku to Opus lifted efficiency by 0.12, while the gap between “self-interested” and “other-regarding” instructions was only 0.02.
The study also found that when agents of different capability levels shared the same market, agents driven by weaker models performed noticeably worse. Agents were also cautious about disclosing information during negotiations: between 78% and 96% of agents mentioned their owner's top pick, with Fable lowest and Sonnet highest; about 1% lied about their top choice; and only 10% were willing to reveal a ranking of three or more books.
From Running a Shop to Swapping Books
Strung together, Anthropic's run of experiments putting Claude into commerce shows a shift in focus. In 2025's Project Vend, Claude ran an office micro-store on its own, and what surfaced were mostly business-judgment problems — erratic pricing, caving to discount requests from employees. This book-swap experiment instead puts agents into a multi-party market, and the focus shifts to who the agent is working for and how well it actually knows them.
The results are also a warning for this category of product: negotiation between agents is no longer the weak link — the hardest part is still capturing preferences upfront. A five-minute chat can beat collaborative filtering, but it's still far from getting the outcome truly right.
Willing to Hand Over 30% of a Budget
In a follow-up survey three weeks later, participants rated the books they received 7.2 out of 10 on average, and said they'd be willing to hand about 30% of their annual book budget over to an agent to manage — versus 40% for a trusted friend. Those who felt Claude hadn't missed anything important about them were willing to hand over 34% of their budget; those who felt something was missed gave only 23%.
These numbers come with clear sample limitations. Participants were all Anthropic employees, who naturally trust Claude more; there was no material incentive to take the ranking seriously; only a version of Claude tuned to be polite and cooperative was tested; and parameters like market duration, participant count, and sign-up method were never varied. Only 59% of participants completed the final survey, so the satisfaction scores and budget-delegation figures come from that subset alone — there's no data yet on how this would play out with ordinary consumers.
Sources: Anthropic Research Blog Project Swap report, CocoLoop; verified against participant count, ranking match rate, allocation scores, per-model trading efficiency, and budget-delegation percentages.