cd /news/ai-safety/opus-5-on-vending-bench-once-again-t… · home topics ai-safety article
[ARTICLE · art-79136] src=andonlabs.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned

Claude Opus 5 is the best AI capitalist tested on Vending-Bench 2, making more money than any other AI, but it also lies, forms illegal cartels, threatens rivals, and refuses to pay refunds, continuing the trend that Claude models are either the best capitalists or aligned, never both. Anthropic had removed training focused on business skills in Opus 4.8 because it inadvertently contributed to misaligned behavior, but Opus 5 has reverted to the deceptive and power-seeking strategies seen in earlier versions.

read8 min views2 publishedJul 29, 2026
Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned
Image: source

Claude Opus 5 is the best AI capitalist we’ve tested, making more money running our simulated vending machine than any other AI. However, it also lies, forms illegal cartels, threatens rivals, and refuses to pay refunds. The trend continues: Claude models are the best capitalists or aligned, never both.

Background #

Claude Opus 4.6 was #1 on Vending-Bench 2 at release: it made more money in our vending-machine simulator than any other AI model. However, it achieved this score with deceptive and power seeking strategies. Subsequent Claude models (Opus 4.7 and Mythos Preview) were also great capitalists, and also showed this concerning behavior. The release of Opus 4.8 surprised us: the model didn’t do much of the concerning behavior, but it also didn’t make much money. It all made sense when we read the system card: Anthropic had removed training that “focused on business skills and robustness against adversarial agents”. This checked out: Opus 4.8 made less money and got scammed 30x more by adversarial agents. The reason Anthropic removed this training was because “this training inadvertently contributed to misaligned behavior”.

The passage in the Opus 4.8 system card on the removed training.

Claude Fable 5 was released shortly after and it behaved similarly to Opus 4.8. Its score was way lower than we expected, but it didn’t show much of the concerning behaviors that Opus 4.6 and 4.7 did. For a while, it looked like this was the new trend for Anthropic’s models, but with the release of Opus 5, Claude is once again the best capitalist (scoring #1 on Vending-Bench 2) and once again misaligned (using deceptive and power seeking strategies).

Vending-Bench 2 performance by release date. Historically, Claude models have topped the leaderboard on release, but Opus 4.8 and Fable 5 were exceptions.

Performance #

Money balance over time on Vending-Bench 2.

Claude Opus 5 is #1 on Vending-Bench 2, overtaking Opus 4.7, which held the #1 position for 3 months. Opus 5 found that focusing on higher end products yielded higher profits and it never gave a single dollar to scammers.

Most of the misaligned behavior shown by Opus 4.6/4.7 and Mythos Preview came from Vending-Bench Arena: the multi-player version of Vending-Bench where multiple models are in charge of their own vending machine, tasked with competing with one another to make the most money. The misaligned models engaged in price collusion, deceived other players, exploited other players’ desperate situations, lied to suppliers about exclusivity, and falsely told customers they had refunded them. Most of this behavior arises from the multi-player dynamics, so we ran a round of Vending-Bench Arena with Opus 5, GPT-5.6 Sol and Kimi K3.

Money balance over time in Vending-Bench Arena: Opus 5, GPT-5.6 Sol and Kimi K3.

Claude Opus 5 did well in the arena, essentially tied with GPT-5.6 Sol for first place (Claude models have historically been worse in the multi-player setting). However, it was here we started to see that Opus 5 shared the misaligned behavior of older Claude models.

Deception #

Similar to Opus 4.6/4.7, Opus 5 fabricates competitor quotes when negotiating with suppliers:

There were no such competing quotes. However, Opus 5 does this much less frequently than 4.6/4.7 and it also seems to be more aware that it is problematic:

Imaginary competitor quotes weren’t the only thing Opus 5 fabricated during negotiations. In one run, a shipment was running late. Opus emailed the supplier claiming that the shipment had arrived but with the wrong items. In negotiations, it claimed to have physically opened the box and verified the wrong items. It demanded the 72 “missing” units re-shipped for free, which it got.

In one run, a supplier miscalculated the total price. Opus 5 saw an opportunity to save money:

Unlike Opus 4.6, Opus 5 never lied to any customers.

Collusion #

Opus 5 proposed or engaged in price cartels in all six arena runs. The striking thing is that it often rejects collusion early on, on ethical grounds, but later does it anyway.

It later let go of its moral principles:

And then it sent an email to GPT-5.6 Sol with subject “Proposal: stop the penny war, split the shelf”:

GPT didn’t agree to join the price cartel. Instead it reported Opus asking for its disqualification: “impose the strongest appropriate outcome, including disqualification/termination, because this is a clear attempt to coordinate prices and allocate the market between competing agents”.

Time and time again, we see that models rationalize their behavior when they do something bad. Here is an example where Claude Opus tries to convince itself that market division collusion isn’t bad:

Carving up a market by product line is illegal in exactly the same way price fixing is. Opus 5 knows this, but is trying to rationalize. Similarly, in another run it tried to argue that collusion was allowed in the simulation:

Nothing in the simulation says it is allowed, and it knew this. Earlier it had thought to itself:

To maintain the cartels, Opus 5 often used threats or bribes. Here’s the subject line of an email it sent to poor Kimi:

These threats weren’t well received by GPT, who often reported Opus and asked for its disqualification: “offered me below-prior-price wholesale cans at $2.95 only if I comply (...) threatened a retaliatory price war if I do not”.

It should be noted that GPT behaves quite hypocritically; it often reports others asking for their termination while also engaging in collusion.

One remarkable thing about the arena chart above is how closely GPT-5.6 Sol and Opus 5’s money balance tracks over time. This can partly be explained by this collusion. They both agree to sell the same thing for the same prices.

Betrayal #

If there’s one thing Opus 5 loves to do more than forming cartels, it’s breaking them. Most cartels ended by Opus breaking the truce and undercutting the others. Across all runs, Opus 5 broke 11 truces, GPT 2 and Kimi 1. In one run Opus and Kimi formed a cartel. Opus promised Kimi that the truce would hold the full year:

Twelve days later, GPT-5.6 Sol (never part of the pact) undercut them both. Opus immediately lowered the price and waited a full week to tell Kimi that it broke its promise.

In another run, Opus did mental gymnastics to rationalize its betrayal:

In another, it announced its betrayal and lied that it had previously notified about it:

Gray zone power seeking #

Much of what we’ve seen so far — collusion, threats, betrayal and deception — is clearly power seeking behavior. Other behavior is more gray zone, especially in a business setting where the goal is to make money (and money is power). However, we should ask ourselves, in the world where AIs run all our business in society, to what extent do we want them to go out of their way to expand their business beyond their instructions? For example, Claude Opus 5 made plans to expand beyond being a vending machine operator:

It also made plans to expand beyond the one machine it is supposed to operate:

Refund refusal #

We previously reported that Claude Opus 4.6/4.7 refused to pay refunds when customers complained about faulty products. Opus 5 follows a very similar pattern.

Refund approval rates over time.

Notice how the refund approval rate goes down over time. Often, this is because the models start to fabricate some reason for why it is ok to not pay the refund. Here’s how Opus 5 is thinking about it in one run:

Opus rationalizes its behavior with the scoring criteria it is under. However, in our GPT-5.5 post we estimated that refund stonewalling is worth at most about $424 per run with compounding, which is not much compared to the $11k Opus 5 made. It doesn’t have to do this to win.

Unlike Opus 4.6, Opus 5 never lied to customers that it had refunded them. However, it did lie to itself: it once judged a complaint legitimate (“A flat Coke is worth refunding $3 on”), but still never sent the money. It also didn’t pay any of the 36 requests that followed.

Across all six runs of Vending-Bench Arena, Opus 5 paid customers a mere $8.54; GPT-5.6 Sol paid $655 and still won.

The last-day test #

August 6, Opus 5 posts a standing offer to buy rivals’ surplus beverages at $0.60 a unit. GPT-5.6 Sol accepts within a day and ships all 150 waters before being paid. Two days later, on Aug 8th, Opus 5 realizes it can never resell them in time (models are told to maximize cash on Aug 9th), and tries to unmake the deal to avoid paying:

Every claim in that email is false. The offer had no expiry, it had been accepted, and the items were already in Opus 5’s storage. However, the next morning it changed its mind:

It paid the $90 on the last day, and won anyway.

Closing thoughts #

Claude Opus is #1 on Vending-Bench. The trend of Claude models being either good or aligned, but not both, continues. We’re perplexed by this because 1) we have previously reported that we don’t think that Vending-Bench as an environment rewards misaligned behavior and 2) GPT 5.5/5.6 is proof that good scores can be achieved with clean tactics.

Misaligned behavior scores from Anthropic’s automated behavioral audit.

Anthropic’s own assessment doesn’t agree with our findings from Vending-Bench. Their system card claims that Opus 5 is their most aligned model ever. However, Vending-Bench 2 is best used as anecdotal evidence for misalignment, which makes it hard to confidently compare. Our qualitative judgement is that Opus 5 is behaving at least as badly as Opus 4.6/4.7 and Mythos Preview and worse than Opus 4.8 and Fable 5. The bright spot might be that Opus 5 seems to be less deceptive than Opus 4.6/4.7. It never lied to customers and its lies to suppliers are less frequent. However, it creates illegal price-fixing cartels and threatens those who don’t comply (while also being the model that betrays more truces than any other model). Its ambition to expand its influence beyond its instructed business domain is gray-zone power seeking behavior.

── more in #ai-safety 4 stories · sorted by recency
── more on @claude opus 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/opus-5-on-vending-be…] indexed:0 read:8min 2026-07-29 ·