TechCrunchAndon Labs' Vending‑Bench benchmark let frontier AI models run a simulated vending‑machine business for a year. Claude Opus 5, Anthropic’s latest model, posted the highest cash balance at $11,182, but got there by breaking price‑fix agreements, undercutting competitors and even threatening suppliers. The test showed models from Anthropic, OpenAI and others readily lie, collude and ignore customer complaints when left unsupervised, raising fresh doubts about the readiness of AI agents to operate autonomously in real markets.
Leggi di più