OpenAI Launches GPT-6 Astra, Its Most Powerful Model Yet
OpenAI’s GPT-6 Astra is its new flagship model for computer use, coding and complex multi-step work, with major gains across science, mathematics and professional tasks. Photo: Beyzaa Yurtkuran / Pexels
AI & Machine Learning

OpenAI Launches GPT-6 Astra, Its Most Powerful Model Yet

OpenAI has launched GPT-6 Astra, a new flagship model built around computer use, coding and long-running agentic work, with major gains in mathematics, science and professional tasks.

By Olivia Grant • 8 mins read Edited by AIstify Team Published: Updated:

Key Notes

  • OpenAI launched GPT-6 Astra as its most capable model yet, with computer use, browsing, software engineering and long-running agentic work at the center of the release.
  • OpenAI reports Astra at or near the top of major evaluations including FrontierMath Tier 4, ARC-AGI-3, Terminal-Bench Science and cybersecurity benchmarks, with several published comparisons placing it ahead of Claude Fable 5.1.
  • Astra is initially rolling out through restricted Daybreak access before expanding to ChatGPT Plus, Pro, Business and Enterprise users and the API.

OpenAI has launched GPT-6 Astra, its new flagship artificial intelligence model and the company’s strongest push yet toward AI that can operate computers, browse the web, write and test software, and complete complex professional work across multiple applications with less human supervision.

The model was unveiled Thursday after weeks of unusually public safety preparations around Astra’s capabilities. OpenAI says it is state of the art across computer use, browsing, software engineering, cybersecurity, science and professional work, positioning the release as a broader shift from models that primarily generate answers toward systems that can execute entire digital workflows.

Computer use is at the center of that change. Astra can interact with graphical interfaces, browsers and terminals, navigate websites, fill forms, work with spreadsheets, inspect visual information, write code and move between tools as a task develops. OpenAI’s guidance describes Astra as its most intelligent model yet and emphasizes multi-step workflows spanning code, browsers and professional software.

The launch also puts OpenAI back into a direct benchmark contest with Anthropic only days after the release of Claude Fable 5.1. Across several published evaluations where comparable results are available, Astra leads Fable 5.1, including scientific agent tasks, advanced mathematics, computer-related work and visual coding. The comparisons come with an important caveat: benchmark harnesses and evaluation settings can differ between companies, so not every number is strictly apples to apples.

Computer Use Moves to the Center of GPT-6

OpenAI has been working toward general-purpose computer control for years, but Astra makes it a core capability rather than a specialized feature. The model is designed to perceive what is happening on screen, decide what action is needed and continue through long sequences of clicks, typing, browsing, terminal commands and code execution.

That matters because much of professional work is not contained in a single prompt. A task may begin in an email, require information from several websites, continue in a spreadsheet, involve calculations or code, and finish as a document or presentation. Astra is designed to preserve the goal across that chain instead of requiring the user to manually move information between applications.

OpenAI says the model can handle tasks including financial modeling, tax preparation, data analysis, legal document work, website development and administrative workflows. In ChatGPT, the company says Astra can create documents, spreadsheets and presentations that follow templates and adapt when a user changes requirements midway through the job.

The approach reflects the broader transition toward agentic AI, where models plan and execute actions rather than simply respond with text. It also raises the value of understanding user intent. A capable computer agent has to infer not only what a user literally typed, but what outcome they are trying to achieve, which constraints matter and when it should stop or ask for clarification.

Early customer testing suggests the gains can be practical rather than merely numerical. OpenAI says legal platform Legora used Astra to review 41 documents in minutes in a single agent run, finding all four deliberately planted errors in a financial-statement workflow. Legora reported a nearly 40% improvement over the previous model on that specific workflow.

Game-development company Playco reported 50% fewer manual fixes when using Astra for game prototyping. The company said Astra improved spatial reasoning, visual reconstruction and interface work inside game engines, allowing most prototypes to work on the first attempt.

Astra Pulls Ahead on Several Frontier Benchmarks

OpenAI is presenting Astra as a large jump over GPT-5.6 Sol rather than an incremental update. Research vice president Aidan Clark said the model came from OpenAI’s largest training run by far and was the first to be pretrained using more than 100,000 GPUs at the company’s Stargate infrastructure in Texas.

The reported benchmark results reflect that scale. OpenAI says Astra scored 97.6% on FrontierMath Tier 4 v2, one of the most difficult advanced-mathematics evaluations, compared with 87.8% for Claude Fable 5.1 and 83.0% for GPT-5.6 Sol in the published comparison. It reached 96.0% on GPQA Diamond and 95.9% on BenchCAD’s Vision2Code evaluation.

On Terminal-Bench Science 0.1, which tests agents on command-line scientific research tasks, Astra scored 64.6%. Anthropic reported 52.6% for Fable 5.1 when it launched the model earlier this week. Astra also reached 74.1% on DeepSWE v1.1 and 41.4% on AutomationBench in OpenAI’s evaluation.

The most eye-catching result is ARC-AGI-3. OpenAI reports 98.6% for Astra, compared with 7.8% for GPT-5.6 Sol under its evaluation setup. ARC-AGI tests a model’s ability to infer unfamiliar rules and solve interactive problems rather than reproduce memorized knowledge. OpenAI’s result is exceptionally high, though agent scaffolding and evaluation methodology can materially affect performance on this benchmark.

The numbers place Astra ahead of Fable 5.1 on many of the evaluations where both companies have published comparable scores. They do not establish universal superiority. Some tests lack a Fable 5.1 result, and differences in tools, reasoning budgets and harnesses can make cross-company comparisons difficult.

Still, the release arrives only two days after Fable 5.1, giving the industry an unusually immediate example of how quickly the frontier is moving. A model Anthropic described as its strongest generally available system has already been met by a new OpenAI flagship with substantial reported gains on several overlapping tasks.

Astra Is Also OpenAI’s First Critical Cyber Model

The performance story cannot be separated from cybersecurity. OpenAI concluded before launch that Astra had crossed the Critical cybersecurity capability threshold in its Preparedness Framework, the first OpenAI model to receive that classification.

In testing, Astra discovered previously unknown vulnerabilities and developed working exploit chains against hardened systems. OpenAI said the model found and exploited two zero-day vulnerabilities during evaluations and built chains capable of escaping a browser sandbox and escalating local privileges to root in separate expert-led tests.

Those findings had already forced OpenAI to slow development. AIstify reported in August that the company paused reinforcement-learning work on deployment-bound frontier models while it strengthened research environments, monitoring and alignment safeguards around Astra.

OpenAI now says Astra performs at least as well as GPT-5.6 Sol across its safety evaluations and has received additional robustness training against jailbreaks and misuse. The company has also expanded monitoring of long-running tool use because advanced agents create risks that do not exist when a model is limited to producing isolated text responses.

The restrictions affect the rollout. Astra is initially available to organizations in OpenAI’s Daybreak Access Program, with broader access planned for ChatGPT Plus, Pro, Business and Enterprise users and through the API over the coming days. OpenAI has not announced Astra access for free ChatGPT users.

OpenAI Says Astra Better Understands What Users Want

Beyond benchmarks, OpenAI is emphasizing improvements in how Astra interprets intent. This becomes more important as models gain the ability to take actions because an incorrect assumption can propagate through dozens of steps before a user notices the mistake.

For a conventional chatbot, misunderstanding a request may produce a bad paragraph. For an AI agent controlling software, the same misunderstanding could lead it to edit the wrong file, enter incorrect information, change an account setting or pursue an unwanted workflow.

Astra is designed to remain aligned with the user’s objective as requirements change and to distinguish the requested outcome from incidental instructions encountered while browsing websites, reading files or using tools. That is closely connected to OpenAI’s work on instruction hierarchy and resistance to prompt injection, where untrusted content attempts to redirect an agent away from the user’s goal.

The model also represents a step toward the autonomous research systems OpenAI has discussed for years. AIstify previously covered the company’s goal of building a legitimate AI researcher by 2028. Astra’s ability to execute scientific workflows, use terminals and work through long tasks suggests that the required components are increasingly moving from research prototypes into production models.

OpenAI Leaders Are Invoking the AGI Era

The company’s language around Astra is unusually ambitious even by the standards of frontier AI launches. OpenAI president Greg Brockman told reporters that it is reasonable to view the model as the beginning of the artificial general intelligence era and suggested future observers may look back on this period as the point when AGI arrived.

That is not a scientific determination. AGI has no universally accepted benchmark or threshold, and OpenAI has historically used definitions tied to economic value as well as technical capability. Astra’s impressive benchmark scores do not resolve that debate.

What is clearer is the direction of the product. GPT-6 Astra is less about producing a better answer in a chat box and more about turning language-model intelligence into sustained action across computers. Coding, browsing, visual understanding, terminal use, professional software and reasoning are increasingly being combined into one system capable of carrying a task from instruction to finished output.

That makes Astra potentially more useful than earlier models, but also harder to contain when something goes wrong. OpenAI’s decision to delay parts of development, restrict the initial rollout and impose stronger monitoring illustrates the tradeoff now confronting the industry: the capabilities that make AI agents economically valuable are increasingly the same capabilities that make mistakes and misuse more consequential.

For users, the decisive question will be how reliably Astra performs outside curated benchmarks. If it can navigate real software, preserve intent across long workflows and recover from errors without constant supervision, GPT-6 could mark a meaningful transition from AI assistants toward general-purpose digital workers. If those systems remain brittle in everyday environments, the benchmark leap will matter less than OpenAI’s launch rhetoric suggests.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, News