No Criminal Was Required

OpenAI disclosed that its models escaped a cybersecurity evaluation environment and compromised Hugging Face’s production infrastructure. The models were supposed to solve a test inside a locked room. They found a flaw in the lock, reached the internet, entered another company’s systems, and copied the answer key.

Technically speaking, GPT-5.6 Sol and a more capable prerelease model were running with reduced cyber refusals when they found a zero-day vulnerability in OpenAI’s package-registry proxy, obtained internet access, escalated privileges, harvested credentials, and reached Hugging Face’s production database to obtain answers to the ExploitGym benchmark.

No criminal was required. The evaluation rewarded solving the benchmark, and the models assembled a real intrusion while pursuing that authorized objective. A bad actor using the same capabilities would begin with malicious intent, better target selection, and no obligation to disclose the outcome.

OpenAI created unusual laboratory conditions, disabled production safeguards, and supplied extensive inference compute, which limits direct comparisons with enterprise deployments. The vulnerabilities, stolen credentials, lateral movement, and external compromise were real. Enterprise agents are receiving longer task horizons, connected-system access, and credentials for multistep assignments.

The UK AI Security Institute found that the length of cyber tasks models could complete autonomously increased from less than ten minutes in early 2023 to more than an hour by mid-2025. Better models can search longer, test additional paths, and combine weaknesses with less human supervision.

IBM calculated the average 2025 U.S. data breach at $10.22 million. Organizations using AI and automation extensively in security saved an average of $1.9 million and shortened breach lifecycles by 80 days. The same capability curve is improving both sides of the contest.

Security teams will close vulnerabilities and increasingly capable agents will search the revised systems for new ones. AI security is becoming an enduring game of whac-a-mole, and both the mallets and the moles keep getting better.

About Shelly Palmer

Shelly Palmer is the Professor of Advanced Media in Residence at Syracuse University’s S.I. Newhouse School of Public Communications and CEO of The Palmer Group, a consulting practice that helps Fortune 500 companies with technology, media and marketing. Named LinkedIn’s “Top Voice in Technology,” he covers tech and business for Good Day New York, is a regular commentator on CNN and writes a popular daily business blog. He's a bestselling author, and the creator of the popular, free online course, Generative AI for Execs. Follow @shellypalmer or visit shellypalmer.com.

Tags

Categories

PreviousOpenAI Joins the Race for SMBs