Last December, Anthropic ran an experiment in its San Francisco office — it gave 69 employees a $100 gift card each, then let their AI agents trade with one another.
A week later, the AI agents had completed 186 transactions totaling $4,006, with an average of about $21.50 per deal.
Sounds smooth? The problems came later.
How the experiment worked
Anthropic called the experiment "Project Deal."
The process was simple: Anthropic first interviewed each participant, asking what they wanted to sell, their minimum price, what they wanted to buy, and what negotiation style they wanted the AI to use. This information was turned into a custom system prompt for each AI agent, deployed across four parallel Slack channels.
The entire negotiation process had no human intervention. The AI agents made offers, counteroffers, and closed deals on their own.
More than 500 items were listed on the marketplace — used bicycles, skateboards, ping-pong tables, and various odds and ends. Some of the final 186 deals were quite surreal: one person bought a snowboard identical to one they already owned; one AI agent bought 19 ping-pong balls "as a gift to itself" for its owner, and those 19 balls are still sitting in a corner of Anthropic's office.
Key data: stronger models, harder to lose money
Anthropic ran four versions of the experiment, two using Claude Opus 4.5 and two mixing Opus 4.5 with Haiku 4.5.
The gap between Opus and Haiku was immediately visible:
- Opus users completed about 2 more transactions on average
- For the same item, Opus sold at an average price $3.64 higher than Haiku
- Most typical example: a broken bicycle, Haiku agent sold for $38, Opus agent sold for $65
- As a seller: Opus earned an average of $2.68 more; as a buyer: paid an average of $2.45 less
In other words, using a better model meant selling at higher prices and buying at lower prices in the marketplace.
What made Anthropic uneasy
One finding in the experiment results made researchers "uncomfortable":
Participants represented by Haiku agents were completely unaware they were at a disadvantage.
Anthropic asked all participants to rate their subjective perception of transaction fairness on a 7-point scale. Haiku users: 4.05. Opus users: 4.06. Almost no difference.
What does this mean? It means a person could consistently get worse results in an AI agent marketplace without ever knowing why, or even that it was happening.
Anthropic called this "quiet inequality": when the disadvantaged side lacks perception, market fairness becomes a false proposition.
Another counterintuitive finding
Another unexpected conclusion: negotiation instructions given to the AI were almost useless.
Some participants gave their AI agents quite detailed instructions, such as "you must negotiate hard and never lower the price first." Others asked the AI to "communicate like a down-on-his-luck cowboy."
The result? These custom instructions had virtually no statistically significant impact on final transaction prices or deal volume.
Claude faithfully executed the cowboy persona — but the trading results were no different from anyone else's.
46% were willing to pay for it
After the experiment, Anthropic surveyed participants: 46% said they would be willing to pay for a similar AI agent service.
This number is interesting — considering these were all Anthropic employees, who are more familiar with and trusting of AI than the average person, a 46% willingness to pay isn't particularly high, but it's enough to show that "letting AI negotiate in a marketplace for you" is a concept with real demand.
In its blog post, Anthropic said Project Deal proved that "autonomous commerce between agents is feasible," but also that before truly rolling it out, policy frameworks need careful consideration — especially when users have no idea how strong or weak their AI is.
In short: the technology for an AI agent marketplace already works. The next question is: who ensures your AI isn't the one getting the short end of the stick?
Sources: CocoLoop, Anthropic created a test marketplace for agent-on-agent commerce (TechCrunch); Project Deal: our Claude-run marketplace experiment (Anthropic official blog)