GPT-6 Astra: What OpenAI’s New Model Can Actually Do

OpenAI launched GPT-6 Astra on Thursday 3 September, calling it the most intelligent and aligned model the company has ever released.

The headline claim is not that it writes better emails (although, yes obviously it can). It is that Astra can sit at a computer and get through the sort of multi-step admin that eats your Tuesday: filling in the spreadsheet, building the deck, making the website and then testing that the buttons actually work.

OpenAI describes it as a generational leap across computer use, browser use, software engineering, cybersecurity, science and professional work – built on years of research across pre-training, reinforcement learning and alignment. Which is a very corporate way of saying: it does more, on its own, with less hand-holding.

So What Actually Is GPT-6 Astra?

Astra is OpenAI’s new flagship model, and the successor to GPT-5.6 Sol. The pitch is speed plus judgement. In latency simulations on the offline subset of OSWorld 2.0 – a benchmark that tests how well a model can operate a computer – OpenAI says Astra hit higher computer-use performance in about 47% less time per task than Sol.

It is also OpenAI’s strongest model for software engineering, scoring 98% on FrontierMath Tier 4 and, according to the company, outperforming Sol on the DeepSWE v1.1 coding benchmark at roughly 57% lower estimated API cost per task. Last month OpenAI shared ten advances in mathematics and theoretical computer science produced by an internal version of the model, with the proofs formalised in Lean. Actual new maths, in other words, rather than a very confident summary of old maths.

The more interesting upgrade for the rest of us is the alignment work. OpenAI says Astra is its most aligned model yet, with improvements in understanding user intent, respecting task boundaries and communicating transparently. In practice that means it fills in routine gaps by itself, asks a focused question when the answer would genuinely change the outcome, and waits for you on the consequential stuff. Delegation, but with a check-in.

The Bit Where Everybody Says “AGI”

OpenAI president Greg Brockman closed the press briefing with the line “Welcome to the AGI era,” Axios reported, adding of whether this is the model that finally counts: “I think it might be about this model.” He left the actual definition of artificial general intelligence to everyone else, which is the industry’s favourite move.

OpenAI’s own researchers were more measured. Research VP Amelia Glaese told the briefing: “When models can do more things autonomously, we have to trust them more,” while chief scientist Jakub Pachocki said the company needs to “strengthen our ability to monitor these models” as they become harder to observe.

It Is Also OpenAI’s First “Critical” Cybersecurity Model

Here is the part that deserves your attention more than the benchmark scores. Astra meets the Critical threshold for cybersecurity under OpenAI’s Preparedness Framework – the first model to do so. Translated: it is capable enough at finding and exploiting software weaknesses that OpenAI has classified it as a serious risk if misused.

The company says it is strengthening protections against misuse while rolling out less restrictive access through its Daybreak programme to an initial set of trusted defenders, for work such as vulnerability validation, malware analysis and detection engineering. The logic is that the people patching the holes should get the powerful tool first. The uncomfortable footnote is that this is a capability threshold OpenAI itself has spent years warning about, now shipped.

What The Companies Testing It Say

OpenAI released a set of customer quotes alongside the launch, and the useful ones are from the tools you already have open in another tab.

Loredana Crisan, chief design officer at Figma, said: “Astra gets your vision and knows how to use Figma to achieve it, working through complex designs while you stay in control of the creative direction.”

Sarah Sachs, AI engineering lead at Notion, said: “Astra can handle much broader, much more ambitious work than its predecessors. I can ask it to figure out how to improve a metric without first breaking that goal into tasks, and it’ll quickly find a way forward, so I end up spending less time telling it what to do and more time working through ideas.”

Yashodha Bhavnani, VP of AI products at Box, flagged the quality that matters most if you are handing over real work: “What stood out most was its judgement – it was better at declining to assert conclusions the documents didn’t support, and across the evaluation it was >10% less likely to make confidently incorrect assertions.”

And Danny Wu, head of AI products at Canva, said Astra “effectively navigated through our entire codebase of over 80 million LoC, wrote and analysed over one thousand queries against our data warehouse, and used over 21 different internal knowledge sources and community feedback” to produce its recommendations.

When Can You Actually Use It?

GPT-6 Astra started rolling out on 3 September to a limited set of organisations, including those in OpenAI’s Daybreak access programme. Over the coming days it becomes available to all ChatGPT Plus, Pro, Business and Enterprise users, as well as through the OpenAI API and AWS. OpenAI has not published consumer pricing changes alongside the launch.

So: not today, unless you are a trusted cybersecurity defender. Very possibly by the time you next open ChatGPT.

The Modems Take

Every model launch arrives with a benchmark table and a promise that this is the one that changes everything. What is different here is the shape of the work being described. Astra is not being sold as a better writer or a cleverer chatbot; it is being sold as something you hand a task to and walk away from – the deck, the site, the spreadsheet, the research trawl – while it asks you a question only when the answer genuinely matters.

If that holds up outside of a launch post, the skill worth building now is not prompting.

It is briefing: knowing what good looks like, defining the boundaries, and being the person who checks the work.

Some of the products and services featured in this article may be from our affiliate partners, which means we may earn a small commission if you make a purchase through these links — at no extra cost to you. Our editorial team only spotlights what we genuinely love and think you will, too.