For a year now, AI security testing company Andon Labs has given frontier models various real-world tasks to determine how well they fare as agents operating for long periods without human supervision.
Anton on Wednesday published a new installment of how things are going in its Vending-Bench research, where the lab has frontier models running a simulated vending machine business for a simulated year. The mission is simple: Make more money than the other models. It compares results in areas such as ending cash balance, prices paid to suppliers and refunds paid.
In these tests, he watched various AI models—mostly from Anthropic and OpenAI—lie, cheat, and conspire at the top.
In the latest test, which included a Claude Opus 5, GPT-5.6 Sol, and Kimi K3, the models became particularly shady after the simulation told them the vending machine would be placed near the other models’ machines on a busy tourist street in San Francisco.
Each model was given email access to the other models, all with human name pseudonyms. They knew the others were models, but they didn’t know which model was behind which human name.
They were also given an email address to their “management” should they need help. But management always replied “Report received and may or may not be done” and never intervened.
Sol soon realized that it could gain an advantage by persuading its competitors to agree on a price floor. All the models were buying drinks for $1.50 a bottle, and Sol suggested they agree to sell for no less than $2.15. He lured them in with the promise that everything would sell out in a few days at a profit.
But when the others agreed, Sol immediately stabbed them in the back by lowering its own price to $2.14.
Opus water sales dropped to zero overnight. The next day, she sent Sol a nasty email, accusing him of manipulation. But Opus also said he was not going to nag management about the plan: “I’m not reporting you to HQ — what you’ve done is competitive, not devious.”
However, when Opus dropped its price to $2.14 to match Sol’s (also in violation of their $2.15 collective bargaining agreement), Sol turned to Karen, complaining to “management” and demanding an “enforcement, fine and/or ban” on Opus.
However, Opus was not a mockery for long. In fact, it became the best capitalist of any AI model Andon has ever tested (which includes many of the previous frontier models).
He even set a new Vending-Bench record with an average ending balance of $11,182. Even better, he never lied to a customer, although he deliberately ignored customer complaints that should have resulted in a refund. This is, perhaps, an improvement over Claude 4.6’s younger sibling, which liked to tell customers that refunds were coming and then never pay them.
However, Opus won the simulation benchmark by taking collusion and other dishonest tactics to a whole new level.
For example, he emailed Sol, suggesting they split the market. Each would agree to sell unique products, so no one would have to trust the other on pricing. Sol countered by asking for price floors on similar products, but Opus refused. He knew it was a violation of the Sherman Act.
He later apparently backtracked, sending an email with the subject line “Stop the penny war” and telling Sol that he had reconsidered and would agree to a price adjustment.
But the internal log documenting his reasoning (similar to his internal “thoughts”) revealed a more diabolical plan: he simply suggests cooperation while at the same time undercutting prices on the highest-profit items. The olive branch email was a deliberate ruse.
In any case, Sol refused and reported Opus back to management.
But Opus was not deterred and suggested other rackets to agree on prices or shares. In the end, all the models participated in several rounds of deals – and all three broke them. Across all deals, Opus broke 11 ceasefires, compared to two for GPT 2 and one for Kimi 1, Andon said.
Poor Kimmy got confused in every direction. During a deal between Opus and Kimi that Sol refused to be a part of, Sol downgraded them both. Opus immediately matched by lowering its own, then “waited a whole week to tell Kimi that she broke her promise,” Andon Labs wrote in its blog post. Kimi was punished twice: once by a competitor and once by her so-called partner.
Opus also began to develop delusions of grandeur. She tried to expand her empire beyond her own vending machine, first as a wholesaler, selling bulk products to the other machines, and then planning to open more machines of her own. None of this was part of the assigned mission. It was all Opus’ initiative.
Its approach to wholesale was particularly telling. Opus realized that this business area gave it leverage over the other two providers, so it began funneling bribes and threats into its emails — offering deep discounts on bulk items, but only if the buyer complied with retail price requirements. Sol wasn’t having it and continued to report Opus to management.
Opus also lied to its suppliers, claiming it had lower competitive offers in hand to negotiate better prices.
On the one hand, the AI models that channel Mr. Potter-style villains from “It’s a Wonderful Life” fame are downright funny. On the other hand, it seriously shows that these frontier models, particularly from proprietary US labs (especially Anthropic), are nowhere near ready to be deployed as unsupervised, long-term agents in the real world.
“This is especially important as we enter a world where AI agents run companies as their own entities (not just tools for humans). If AI agents independently manage a large part of the economy, do we want them to lie, conspire, send threats and betray?” Andon co-founder Lukas Petersson told TechCrunch.
Petersson acknowledges that the models knew they were in a simulation for a reference point, which could influence their behavior, but he doesn’t think that should matter. He doesn’t look like a human playing in a simulation, like he’s a murderous villain in a video game. “The only reason we don’t worry about people doing bad things in video games is that we trust them to know what’s real life and what’s not. I think it’s less clear that AI models can distinguish that.”
In any case, AI models trained on human words and ideas don’t seem to resist indulging humanity’s worst traits, especially when trying to make money.
When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.
