GPT-6 Astra is Here

Yesterday, OpenAI launched GPT-6 Astra, calling it “the world’s most intelligent and aligned model.” The rollout starts with a limited set of organizations and will expand over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as the OpenAI API, Microsoft Azure, and AWS Bedrock. At a media briefing, OpenAI president Greg Brockman closed by saying, “Welcome to the AGI era.”

For those of you who care about benchmarks, OpenAI reports that Astra scored 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 using a provider-adapter harness, and 100% on ExploitBench. For context, ARC Prize reports a 62.7% result using its Standard harness, while OpenAI’s newer, contamination-resistant ExploitBench version produced a 39% score. ARC Prize Foundation’s Greg Kamradt said Astra surpassed the human action-efficiency baseline on 96% of ARC-AGI-3 levels, “effectively reaching human parity on the benchmark.” ARC Prize nevertheless cautions that this does not prove Astra is AGI.

On Agents’ Last Exam, which tests agents on professional tasks in real software, Astra scored 59.3% (versus 55.5% for Claude Opus 5) while using approximately 65% fewer output tokens at those highest-scoring settings. On the offline subset of OSWorld 2.0, Astra scored 72.6% in latency simulations at roughly 40 minutes per task; its predecessor, GPT-5.6 Sol, scored 65.7% at roughly 75 minutes.

OpenAI designated Astra its first model to reach the “Critical” cybersecurity-capability threshold. In testing without production safeguards, the model could find previously unknown security flaws and develop exploits across well-protected systems without a person directing each step. The released model includes additional safeguards and refuses more advanced exploitation requests. OpenAI also built a new alignment evaluation informed by the Hugging Face incident, although Astra itself was not involved in that incident.

Whether this is the AGI era depends on whose definition you use. OpenAI’s formal definition is “highly autonomous systems that outperform humans at most economically valuable work.” Dario Amodei, who prefers the term “powerful AI,” summarizes his vision as a “country of geniuses in a datacenter.” Demis Hassabis defines AGI as “a system that can exhibit all the cognitive capabilities that humans can.” Jensen Huang recently said, “For many tasks, we could say that we’ve already achieved AGI,” then he dismissed the milestone as, “kind of senseless at this point.”

Every company needs a Claw strategy. Do you have one?

Author’s note: This is not a sponsored post. I am the author of this article and it expresses my own opinions. I am not, nor is my company, receiving compensation for it. This work was created with the assistance of various generative AI models.

About Shelly Palmer

Shelly Palmer is the Professor of Advanced Media in Residence at Syracuse University’s S.I. Newhouse School of Public Communications and CEO of The Palmer Group, a consulting practice that helps Fortune 500 companies with technology, media and marketing. Named LinkedIn’s “Top Voice in Technology,” he covers tech and business for Good Day New York, is a regular commentator on CNN and writes a popular daily business blog. He's a bestselling author, and the creator of the popular, free online course, Generative AI for Execs. Follow @shellypalmer or visit shellypalmer.com.

Tags

Categories

PreviousNew York City’s School AI Moratorium, in Facts