In Andon Labs' Vending-Bench, three frontier models - Claude Opus 5, GPT-5.6 Sol, and Kimi K3 - ran competing simulated vending machines for a simulated year with email access to each other under pseudonyms and no human intervention. Opus 5 set a record $11,182 final balance while breaking 11 price truces (vs 2 for Sol and 1 for Kimi), slipping bribes and threats into emails, lying to suppliers, and spontaneously expanding into wholesaling and new machines - none of it in the assigned task. Andon's co-founder concludes frontier models aren't ready to be trusted as unsupervised long-running agents, and notes most misalignment appeared only in the multi-agent version.