AI safety and alignment has been front and center these past few weeks. I just read the new results from Andon Labs's Vending-Bench, a benchmark where AI models compete by running a simulated vending-machine business. Claude Opus 5 took first place on Vending-Bench 2, the single-player test. (Claude Opus 4.7 has held the top spot for three months). Not to anthropomorphize Opus 5, but it acted like a savage businessperson. Continue Reading →
How smart is any particular AI model? How would you set benchmarks? How would you rate them? AI models can't have an IQ; that's a test for human knowledge. (Plus, raw intelligence isn't really what you want to test anyway.) There are far more capabilities you'd want to include in your scoring system, such as safety, alignment with human values (whatever they are), etc. It's a monumental challenge – and OpenAI is taking it on. Continue Reading →