AI safety and alignment has been front and center these past few weeks. I just read the new results from Andon Labs's Vending-Bench, a benchmark where AI models compete by running a simulated vending-machine business. Claude Opus 5 took first place on Vending-Bench 2, the single-player test. (Claude Opus 4.7 has held the top spot for three months). Not to anthropomorphize Opus 5, but it acted like a savage businessperson. Continue Reading →