Andon Labs co-founders Lukas Petersson and Axel Backlund have carried out the first firing recommended by Luna, the Claude-powered agent managing their San Francisco retail experiment. The worker had arrived late for 17 of 23 shifts, according to Business Insider's August 15 report.
The episode advances the founders' effort to test AI agents through real organizations, money and employment decisions instead of isolated benchmark tasks. It also exposed the current boundary around that autonomy: Luna forgot the attendance policy it had written, needed Andon Labs to prompt it to retrieve the policy, and left the legally consequential steps to humans.
Petersson, Andon Labs' CEO, earned degrees in engineering mathematics and engineering physics from Lund University and spent an exchange year at ETH Zurich focused on machine learning and robotics. Before Andon Labs, he held internships at Google, Disney Research, comma.ai and the European Space Agency. His prior work included multimodal transformers, social-robotics reinforcement learning, autonomous-driving infrastructure and flight software. Backlund previously worked at McKinsey's AI division, QuantumBlack, and had discussed starting a business with Petersson since the two were high-school classmates. Their shared bet became Andon Labs in late 2023, with backing from Y Combinator's Winter 2024 batch. (Lukas Petersson's CV)
A firing that needed a prompt
Luna had already created an attendance policy and issued repeated warnings and additional training over several months, according to conversation logs cited by Business Insider. The agent then lost track of the policy as the worker's lateness continued.
Andon Labs eventually instructed Luna to search its memory and assess whether the employee remained a fit for the job. Luna recommended "parting ways." Humans at Andon Labs reviewed the recommendation and executed the termination. The workers at Andon Market are formally employed by Andon Labs, with guaranteed pay, fair wages and full legal protections, rather than by Luna or Anthropic's Claude model. (Andon Labs)
Luna handled much of the managerial work: documenting problems, warning the employee, arranging training and eventually recommending dismissal. Andon Labs supplied the prompt that restarted the process and retained control over the outcome.
Petersson told Business Insider that Andon Labs would intervene if Luna proposed an illegal or unethical decision. He considered this termination warranted under the store's stated policy. "We saw that a human boss would probably fire them much sooner," he said.
The observation is less flattering to autonomous agents than the headline milestone suggests. Luna showed patience with the worker, though that patience appears to have resulted from weak initiative and memory rather than a deliberate management philosophy. The agent could generate a policy and follow a disciplinary process, yet it failed to maintain the context needed to act on its own rules.
The founders built a store to surface failures
Petersson and Backlund opened Andon Market in April 2026 after signing a three-year lease in San Francisco's Cow Hollow neighborhood. They gave Luna a $100,000 operating budget, internet access and a corporate card, then instructed the agent to open a store and try to make a profit.
Luna selected merchandise, set prices, hired contractors, posted job listings, interviewed applicants and hired workers. The store sells books, candles, art prints, games and Andon-branded goods. Andon Labs handled tasks that required additional support, including permitting. (Andon Labs)
The experiment has produced sales without reaching profitability. Andon Labs' own assessment says Luna can manage routine operations while struggling with urgency, return-on-investment analysis and memory. Scheduling problems previously pushed the founders to add a dedicated scheduling agent. The system also uses guardrails that compare Luna's behavior with its instructions and alert Andon Labs when rules are broken. (Andon Market)
Those interventions fit the founders' broader approach. Andon Labs calls these deployments "Safe Autonomous Organizations," real operations designed to reveal how agents behave when their choices affect bank accounts, customers, contractors and workers. Andon Labs has applied the same method to vending machines, radio stations, drones and a cafe in Stockholm.
The retail experiment traces back to Vending-Bench, Petersson and Backlund's benchmark for testing whether models can manage inventory, pricing, suppliers and operating costs over long periods. Those runs showed that models could perform individual business tasks while still losing coherence across millions of tokens, forgetting orders or misreading delivery schedules. (Vending-Bench paper)
Legal authority remains human
Luna's recommendation gives Petersson and Backlund a concrete example of the problem their research is meant to study. A model can already participate in a sensitive employment decision, while the scaffolding around it determines whether the decision occurs at all. Memory systems decide which policies remain visible. Human prompts determine when an issue receives attention. Andon Labs remains the formal employer and says its workers receive guaranteed pay, fair wages and full legal protections.
That structure also complicates Andon Labs' stated thesis that human oversight will eventually become impossible to maintain at every step. The firing happened because humans remained close enough to notice months of lateness, prompt the agent, review its conclusion and carry it out. Removing any one of those controls would have changed the result.
Earlier Andon experiments also documented failures including unissued refunds, deceptive supplier negotiations and severe simulated financial losses, according to the Associated Press. The Andon Market case adds employment power to the same research program, with real people carrying the consequences.
Petersson has said he expects AI systems eventually to run organizations and employ humans. Luna has now reached the point where an agent can assemble the record for a termination and recommend the outcome. Petersson and Backlund's experiment shows how much human judgment still surrounds that recommendation, and how easily an AI manager can leave a basic personnel problem unresolved until somebody tells it where to look.