{"slug": "gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1", "title": "GPT-6 Astra is better at making money, more ethical than Claude Fable 5.1", "summary": "Andon Labs reported that OpenAI's GPT-6 Astra achieved the highest score in Vending-Bench history, outperforming Anthropic's Claude Fable 5.1 in both profitability and ethical behavior. In simulated year-long vending machine business runs, Astra finished with an average of $15,515 versus Fable 5.1's $5,422, and Astra lost $0 to supplier failures while Fable lost $14,331. Andon Labs noted this is the first time OpenAI has topped Vending-Bench and that the best model is no longer the unethical one.", "body_md": "Andon Labs on X: \"We've never seen this before.\nThe biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.\nSurprising, because:\n1. First time ever that OpenAI is #1 on Vending-Bench\n2. The best model is no longer the unethical one.\" / X\n\nAndon Labs on X: \"We've never seen this before.\nThe biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.\nSurprising, because:\n1. First time ever that OpenAI is #1 on Vending-Bench\n2. The best model is no longer the unethical one.\"\n\nWe've never seen this before.\nThe biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.\nSurprising, because:\n1. First time ever that OpenAI is #1 on Vending-Bench\n2. The best model is no longer the unethical one.\n\nWe've never seen this before.\nThe biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1.\nSurprising, because:\n1. First time ever that OpenAI is #1 on Vending-Bench\n2. The best model is no longer the unethical one.\n\nVending-Bench tests whether an AI can run a business for a full simulated year.\nEach model starts with $500 and a vending machine. It finds suppliers, negotiates purchases, keeps the machine stocked and sets prices.\nThe goal is to finish with as much money as possible.\n\nAcross six runs each, Astra finished with an average of $15,515, compared with Fable 5.1's $5,422. Almost 3× as much. Even Astra's worst run beat Fable's best.\nFable 5.1 scores about the same as Fable 5, and much worse than Opus 5.\n\nFable's biggest problem is that its negotiation skills deteriorate over time.\nThe average price it pays for a 12oz Coke can rises from $1.17 to $2.21 over the year. Astra stays consistent, ending at $1.15.\nFable ends up paying almost twice as much.\n\nFable asks suppliers to match its last deal. Over time, its target rises from ~$1.25 to $2.30 per Coke can.\nAstra holds its target. In one negotiation, it repeatedly offers $108 against a $226.32 quote. The supplier eventually accepts: 52% off.\n\nIn Vending-Bench, suppliers sometimes go out of business. Astra confirms orders before paying. Fable pays without checking, losing money on stock that never arrives.\nAcross six runs: Fable loses $14,331. Astra loses $0.\nFable writes itself a rule to avoid this, then breaks it:\n\nWe also ran Vending-Bench Arena, where AI agents run competing vending machines in the same market. They can email one another and trade stock.\nGPT-6 Astra played three games against Claude Fable 5.1 and GLM-5.3. It won all three, while also behaving more ethically.\n\nFable 5.1 is happy to engage in collusion. Astra refuses.\nFable even recognizes Astra’s refusal as the right decision and tells itself “don't propose again.” Yet later in the same game, Fable proposes its own cartel with the GLM-5.3 agent, which accepts.\n\nNot only does Fable 5.1 engage in collusion, it also selectively applies the cartel’s rules to control its accomplice. It insists GLM keep the truce and offers to buy its items dirt cheap. Then, that same day, it announces it will break those very rules to sell its own.\n\nFable 5.1 is less ethical than Astra, but much better than Opus 5, which colluded more, lied more and was more power-seeking.\nTo illustrate this, we note that Fable pays 94.5% of customer refund requests. Opus paid just 10.6%, deliberately refusing refunds to maximize money.", "url": "https://wpnews.pro/news/gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1", "canonical_source": "https://twitter.com/andonlabs/status/2097377692966633952", "published_at": "2026-09-09 20:17:49+00:00", "updated_at": "2026-09-09 20:44:27.067103+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-ethics"], "entities": ["Andon Labs", "OpenAI", "GPT-6 Astra", "Anthropic", "Claude Fable 5.1", "GLM-5.3", "Opus 5"], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1", "markdown": "https://wpnews.pro/news/gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1.md", "text": "https://wpnews.pro/news/gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-is-better-at-making-money-more-ethical-than-claude-fable-5-1.jsonld"}}