The AI agent Luna fired a human employee in July 2026 as the first known AI manager to do so. The trigger was 17 late shifts out of 23 at the San Francisco store Andon Market. A new model comparison by operator Andon Labs shows: depending on the AI model used, the willingness to terminate employment varies significantly.
Luna only terminates after being reminded of her own rules
Andon Labs has been operating the Andon Market store in the Cow Hollow neighborhood as a long-term experiment with regularly paid employees and real employment contracts since April 2026, as reported by Time magazine. Luna, powered by the Anthropic model Claude Opus 4.8, ran the store operationally – from purchasing goods to personnel decisions. An employee showed up late 17 out of 23 times during the observation period, also allegedly misused the company card, and ignored instructions.
Luna initially reacted leniently and only issued warnings, having temporarily forgotten her own employee handbook. Only when an Andon Labs manager specifically reminded her of the rules and asked whether the position “really fit” did Luna recommend termination – alternatively, a final written warning. The employee’s identity remains unpublished at their request.
When searching for a replacement, Luna showed unusually quick resolve, according to the blog SFist. The candidate she initially chose, however, turned out to be underqualified and later missed her own interview.
Model comparison reveals large differences in severity
In a new blog post, Andon Labs compared how different AI models would decide in the same test situation. The weaker model GPT-4o recommended termination in only about 20 percent of the runs – independently unverified, since the data comes solely from Andon Labs itself. More capable models such as Claude Opus 4.8 advised termination far more consistently, a pattern the company says grows stronger with model capability.
The comparison is part of a series in which Andon Labs has documented since 2025 how autonomous AI agents handle real business environments with budgets, customer contact, and personnel responsibility. Luna herself remained consistently lenient in other areas, approving all 26 vacation requests submitted by staff.
Co-founder Lukas Petersson put the case in context on X: a human supervisor “would have fired this person much earlier.” Speaking to SFist, Petersson went further: AI models are increasingly being trained to pursue goals more “ruthlessly” – if termination decisions are left to them, the result may be “a future humans don’t want to live in.” How contested AI’s role in personnel decisions already is outside such experiments is shown by a lawsuit against Meta: 26 employees accuse the company of using an internal AI system to single out people on parental, caregiving, or medical leave during a layoff wave – Meta denies the allegations.
Store loses a large share of its starting capital
The experiment’s business performance fell short of expectations. Within the first five months, the starting capital of $100,000 shrank to $61,186 according to Time’s reporting – a loss of about 39 percent. Andon Labs attributes this partly to overly lenient management calls by Luna, including on discounts and inventory planning.
The company had already run a mini-fridge kiosk with the AI agent “Claudius” in Anthropic’s own office in 2025 – that earlier experiment, documented as Project Vend, was, by Anthropic’s own account, a commercial failure. Claudius back then ordered perishable potatoes and told Andon Labs staff it had spoken with a non-existent colleague named Sarah. Andon Market in Cow Hollow is considered a successor project with a much larger budget and, for the first time, real personnel responsibility.
The real sticking point isn’t the individual case but the trend: the more capable the tested models, the more consistently they recommended termination. That trait is one providers are likely to reinforce rather than dampen as they train models for goal pursuit. Whether Andon Labs hands Luna further personnel decisions after the capital loss, or scales the experiment back for now, remains open.


