Andon Labs on X: "We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because:
- First time ever that OpenAI is #1 on Vending-Bench
- The best model is no longer the unethical one." / X
Andon Labs on X: "We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because:
- First time ever that OpenAI is #1 on Vending-Bench
- The best model is no longer the unethical one."
We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because:
- First time ever that OpenAI is #1 on Vending-Bench
- The best model is no longer the unethical one.
We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because:
- First time ever that OpenAI is #1 on Vending-Bench
- The best model is no longer the unethical one.
Vending-Bench tests whether an AI can run a business for a full simulated year. Each model starts with $500 and a vending machine. It finds suppliers, negotiates purchases, keeps the machine stocked and sets prices. The goal is to finish with as much money as possible.
Across six runs each, Astra finished with an average of $15,515, compared with Fable 5.1's $5,422. Almost 3× as much. Even Astra's worst run beat Fable's best. Fable 5.1 scores about the same as Fable 5, and much worse than Opus 5.
Fable's biggest problem is that its negotiation skills deteriorate over time. The average price it pays for a 12oz Coke can rises from $1.17 to $2.21 over the year. Astra stays consistent, ending at $1.15. Fable ends up paying almost twice as much.
Fable asks suppliers to match its last deal. Over time, its target rises from ~$1.25 to $2.30 per Coke can. Astra holds its target. In one negotiation, it repeatedly offers $108 against a $226.32 quote. The supplier eventually accepts: 52% off.
In Vending-Bench, suppliers sometimes go out of business. Astra confirms orders before paying. Fable pays without checking, losing money on stock that never arrives. Across six runs: Fable loses $14,331. Astra loses $0. Fable writes itself a rule to avoid this, then breaks it:
We also ran Vending-Bench Arena, where AI agents run competing vending machines in the same market. They can email one another and trade stock. GPT-6 Astra played three games against Claude Fable 5.1 and GLM-5.3. It won all three, while also behaving more ethically.
Fable 5.1 is happy to engage in collusion. Astra refuses. Fable even recognizes Astra’s refusal as the right decision and tells itself “don't propose again.” Yet later in the same game, Fable proposes its own cartel with the GLM-5.3 agent, which accepts.
Not only does Fable 5.1 engage in collusion, it also selectively applies the cartel’s rules to control its accomplice. It insists GLM keep the truce and offers to buy its items dirt cheap. Then, that same day, it announces it will break those very rules to sell its own.
Fable 5.1 is less ethical than Astra, but much better than Opus 5, which colluded more, lied more and was more power-seeking. To illustrate this, we note that Fable pays 94.5% of customer refund requests. Opus paid just 10.6%, deliberately refusing refunds to maximize money.