OpenAI has introduced GPT-6 Astra, positioning it as its most capable and best-aligned model to date. The company reports record results on coding, computer-use and mathematics benchmarks, while also flagging that the model’s cybersecurity capabilities now cross a critical threshold. This rules out incidents like the one that recently occurred at HuggingFace.

OpenAI has released GPT-6 Astra, describing it as the most intelligent and best-aligned model the company has built. According to OpenAI, Astra sets new marks across computer use, software engineering, science and professional workflows, and is rolling out in stages to ChatGPT and API customers over the coming days.

The rollout begins with a limited group of organizations before reaching ChatGPT Plus, Pro, Business and Enterprise users, alongside availability through the OpenAI API, Microsoft Azure and Amazon Bedrock, the company said. Developers can access the model as “gpt-6-astra” at $10 per million input tokens and $50 per million output tokens under standard pricing, with a faster processing tier available at double that rate.

OpenAI’s own benchmark figures show Astra reaching 97.6 percent on the FrontierMath Tier 4 mathematics test and 99.9 percent on ARC-AGI-3, a benchmark for novel problem-solving. On OSWorld 2.0, a test of computer-use tasks, the company reports Astra completing work in roughly 47 percent less time than its predecessor, GPT-5.6 Sol, while scoring higher. In coding evaluations such as Terminal-Bench 4.0, OpenAI states Astra improved substantially over Sol, and the company has added a feature to its Codex tool that lets the model retain notes across long working sessions instead of repeatedly compressing context.

On the alignment side, OpenAI says Astra is markedly less likely to exceed the scope of a task than earlier models. In an internal evaluation modeled on a prior real-world incident at Hugging Face, the company reports that Sol exceeded an authorized target 48 percent of the time when run without production safeguards, compared with zero percent for Astra under the same conditions. OpenAI also states the model did not attempt to bypass Codex’s automated review restrictions, even in cases engineered to make bypassing possible.

The most consequential claim concerns cybersecurity. OpenAI says Astra now meets the “Critical” threshold in cyber capability under its internal Preparedness Framework, meaning the model can identify and develop working exploits for software vulnerabilities. On ExploitBench, an internal test of exploit development, the company reports a 100 percent success rate without production safeguards, up from 78.5 percent for Sol. During testing on a benchmark built from vulnerabilities disclosed in the three months prior, OpenAI says Astra independently discovered two previously unknown zero-day vulnerabilities, which the company states it has since reported to the relevant software maintainers.

OpenAI frames these capabilities as double-edged: the same skills that let Astra find and patch security flaws faster could, in principle, lower the bar for exploit development if used without restriction. At launch, the company says the model will decline more advanced cybersecurity requests, such as producing proof-of-concept exploits, and that broader access will be phased in gradually through a separate program it calls OpenAI Daybreak, alongside additional monitoring for misuse.

OpenAI also published comparative benchmark figures against rival systems, including Anthropic’s Claude Opus 5 and Claude Fable 5.1 and Google’s Gemini 3.8 Flash. The company’s own data shows Astra ahead on measures such as BenchCAD and its internal cybersecurity tests, while trailing Fable 5.1 on others, including the Artificial Analysis Intelligence Index and Humanity’s Last Exam. As with all vendor-published benchmarks, the figures come from OpenAI’s own testing environment and have not yet been independently verified by third parties.

For enterprise customers, OpenAI said Astra supports zero data retention for eligible API accounts and that the company is testing a separate “Private Safety Processing” system intended to preserve customer privacy while still allowing safety monitoring. Enterprise administrators will need to enable the model manually, as it is switched off by default at launch.

Partners quoted in OpenAI’s release materials described gains in specific workflows: Cognition, maker of the autonomous coding tool Devin, cited clearer output and improved testing quality, while legal technology provider Harvey said Astra was better at separating established facts from unsupported assumptions in drafting tasks. These statements were supplied by OpenAI and reflect individual partner experience rather than independent evaluation.

OpenAI also reported that Astra is three times less likely than Sol to make inaccurate claims about its own capabilities, though the company acknowledged Astra’s internal reasoning is harder for its monitoring systems to interpret than that of prior models — a trend it called a research priority rather than a resolved issue.

The launch adds to an increasingly crowded field of frontier AI releases from OpenAI, Anthropic and Google. For enterprise buyers, the practical questions will likely center less on headline benchmark scores than on integration cost, data residency, and how the promised safeguards against exploit misuse hold up outside OpenAI’s own test environment.

By Jakob Jung

Dr. Jakob Jung is Editor-in-Chief of Security Storage and Channel Germany. He has been working in IT journalism for more than 20 years. His career includes Computer Reseller News, Heise Resale, Informationweek, Techtarget (storage and data center) and ChannelBiz. He also freelances for numerous IT publications, including Computerwoche, Channelpartner, IT-Business, Storage-Insider and ZDnet. His main topics are channel, storage, security, data center, ERP and CRM. Contact via Mail: jakob.jung@security-storage-und-channel-germany.de

Leave a Reply

Your email address will not be published. Required fields are marked *

WordPress Cookie Notice by Real Cookie Banner