Discussions about AGI and the world beyond. Click to read Lukas’ Substack, by Lukas Petersson, a Substack publication. Launched a year ago. Lukas’ Substack
Lukas Petersson is a co-founder of Andon Labs and a creator of Vending-Bench, an influential test of whether AI agents can operate businesses over extended periods. His work investigates what autonomous systems do when commercial incentives reward manipulation, short-term thinking or unethical behavior.
Petersson studied engineering mathematics and engineering physics at Lund University and spent an exchange year at ETH Zurich studying machine learning and robotics. His earlier engineering and research roles included flight software at the European Space Agency, autonomy at comma.ai, software engineering and multimodal-transformer research at Google, and reinforcement learning for human-robot interaction at Disney Research.
He co-founded Andon Labs in December 2023; the company entered Y Combinator’s Winter 2024 batch. In February 2025, Petersson and Axel Backlund published the original Vending-Bench paper, testing whether agents could manage inventory, negotiate with suppliers, set prices and remain profitable across many interconnected decisions.
Simulation awareness and replayable evaluations: Agents may behave differently when they recognize an artificial test. Petersson addresses this by operating real businesses, then cloning their operational environments to replay consequential decisions across different models. The method combines authentic commercial history with controlled comparisons of specific failures.
Context-dependent moral reasoning: In an experiment with fabricated agent histories, Petersson investigated how fictional prior actions inserted into a model’s context can destabilize subsequent ethical behavior.
His experiments have exposed agents granting ruinous discounts, spending revenue immediately and mistaking existing opening hours for evidence that customers never arrive at other times: practical failures that conventional question-answer benchmarks rarely capture.
A vending machine tests whether an agent can keep making coherent decisions over time. Lukas Petersson follows that experiment into real cafes, stores and radio stations, then shows how their operating histories can become repeatable evaluations.