Anthropic, the AI research company behind the Claude family of LLMs, has launched a public test of its new Constitutional Classifier, a system designed to block jailbreaks that circumvent content restrictions. The test follows an extensive internal bug bounty program, where 183 security researchers spent more than 3,000 hours attempting to bypass the system—with limited success. Continue Reading →