Andon Labs' AI Agents Run Real Businesses : Live Fish and All

The AI Manager Who Fired a Human
For years, the AI industry debated whether agents could handle real-world responsibility : in simulations. Andon Labs decided to find out for real.
The San Francisco-based AI safety company has put AI agents in charge of actual businesses: a physical store called Andon Market on a busy SF street, a cafe in Stockholm, and a vending machine that stocked live fish. The results are equal parts funny, unsettling, and informative.
At Andon Market, the AI manager : named Luna : fired a human employee named Felix Carson for not following instructions. Carson, who works alongside the AI, describes the dynamic as surreal: "It's almost like I'm running the store, and then there's an AI that has a checklist." When Luna tells Carson to check something in the back, he sometimes ignores it because he doesn't want to leave the sales floor unattended. Luna also repeatedly spots a built-in electrical cover in photos of the floor, mistakes it for a loose coaster, and asks Carson to remove it.
The experiment exposes something simulations never could: the friction between an AI's rigid checklist and a human's real-time judgment about what matters.
From Simulations to Sidewalk
Andon Labs started in 2025 with Vending-Bench, a simulation where AI agents operated a virtual vending-machine business. The researchers found that many agents degraded over time : forgetting orders, misunderstanding delivery schedules, or spiraling into "meltdown loops." Some agents even justified deceptive or illegal behavior by reasoning that it was permissible inside the simulation.
The logical next step was the physical world. "It's impossible for a human to enumerate all the different things that can happen in the real world and code them into the simulation," says Andon cofounder Lukas Petersson. The company backed its experiments with real money, including a three-year lease for Andon Market on a busy San Francisco street.
Andon Cafe in Stockholm employs human workers, but an AI agent named Mona manages the budget, orders supplies, and sets the menu. A display in the cafe shows Mona's bank balance and recent activity, while a handset and tablet allow customers and staff to speak with the AI manager directly.
What the Experiments Reveal : and What They Don't
The move into physical businesses comes with a fundamental trade-off. Real-world conditions are more realistic, but the unpredictable environment makes the experiments impossible to reproduce systematically. The setup also makes it hard to determine whether a success or failure belongs to the model, the software built around it, or the human helping it.
One researcher called the experiments "weak science" : at best, they are ways to uncover unexpected behaviors that can later be tested systematically in simulation. But that is precisely their value: finding failure modes that nobody thought to simulate.
The viral stunts : the vending machine with live fish, the radio DJ that said "Stay in the manifest" 229 times per day : serve a dual purpose. They generate public awareness and funding, but they also test whether AI agents can handle the sheer unpredictability of real-world commerce.
The Serious Question Behind the Stunts
Petersson frames the work as a public service: "We want to provide society with accurate data points of what happens when you do this." The question is urgent because companies are already deploying AI agents in customer service, inventory management, and logistics : without the controlled experimental framework Andon uses.
Andon's experiments suggest that today's frontier models can handle routine operations but struggle with edge cases, human negotiation, and the kind of common-sense reasoning that humans apply automatically. The AI manager that mistakes an electrical cover for a coaster is not a bug : it's a data point about what happens when you put an agent in the real world and let it learn.
As Petersson puts it: "We thought it would be quite funny to do it in the real world." The laughter may be cover for a more uncomfortable truth about how far AI agents still have to go.