GPT-6 Astra: OpenAI's Next Frontier Model

GPT-6 Astra introduces major breakthroughs in computer use, software engineering, science, and alignment.

24 cards · 3 min · tap to begin

From · · 3 min

GPT-6 Astra: OpenAI's Next Frontier Model

GPT-6 Astra introduces major breakthroughs in computer use, software engineering, science, and alignment.

In brief

GPT-6 Astra introduces major breakthroughs in computer use, software engineering, science, and alignment. GPT-6 Astra pairs unprecedented computer use and coding performance with rigorous alignment and real-time safety monitoring. Originally reported by OpenAI.

Introducing GPT-6 Astra

GPT-6 Astra is OpenAI's most intelligent and aligned model to date, combining years of research in pre-training, reinforcement learning, and alignment.

We’re introducing GPT‑6 Astra, the world’s most intelligent and aligned model.

Saturating frontier benchmarks

The model sets new benchmark highs, reaching 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.

Not only is this the best model we’ve ever tested, but it also represents a meaningful step change in frontier-model performance

Cheaper scientific workflows

On Terminal-Bench Science 0.1, Astra completed research workflows with a 64.6% success rate while reducing estimated API costs by roughly 31%.

Unprecedented scope discipline

When faced with impossible tasks, Astra never went beyond its authorized target, compared to previous models that exceeded scope 48% of the time.

Next-level computer control

Astra autonomously operates desktop applications to fill out forms, update CRMs, run web research, execute software installs, and conduct QA tests.

Faster knowledge work execution

The model scored 59.3% on Agents' Last Exam while using 65% fewer output tokens, completing OSWorld tasks in 47% less time than previous generations.

Accelerating physical hardware design

Astra can perform complex printed circuit board layouts in KiCad, placing components and routing copper connections to accelerate hardware engineering.

Optimized harness for speed

Paired with an updated Codex harness, Astra completes web and desktop automation tasks 1.9 times faster on the Mind2Web benchmark.

Professional document creation

The model adheres strictly to enterprise templates, generating polished slide decks, spreadsheets, and reports that match existing visual and writing styles.

Precision 3D CAD modeling

On BenchCAD, Astra reconstructed 3D objects from multi-view renders with a 95.9% geometric overlap score at up to 86% lower API costs.

Generating interactive 3D worlds

Astra can model architectural assets in Blender and assemble them into walkable Unreal Engine 5 scenes or full WebGL browser games from text prompts.

Smart ambiguity resolution

When instructions lack detail, Astra asks targeted questions asynchronously while continuing unblocked work, making sensible defaults when unanswered.

Staying oriented through feedback

Unlike older models that lost track of original goals when receiving mid-task feedback, Astra adapts to new requirements without dropping prior constraints.

State of the art in coding

Astra hit 57.9% on Terminal-Bench 4.0, outperforming rival models on complex terminal engineering tasks with significantly lower token consumption.

produces code that requires less iteration to reach production quality

Persistent notes across long sessions

Codex now allows Astra to keep structured notes across context windows instead of relying on destructive compression summaries that erase vital details.

Solving open math problems

Astra helped prove that infinitely many prime pairs occur within a distance of 186, while also breaking an 80-year-old bound on large prime gaps.

The story is: end of one era, start of another.

Navigating complex scientific tools

Hitting 96.0% on GPQA Diamond, Astra can operate life-science software directly to inspect genetic sequencing quality and visualize data variations.

Reaching critical cyber thresholds

Astra triggers the Critical threshold under OpenAI's Preparedness Framework, scoring 100% on ExploitBench and 42.4% on ExploitGym.

Uncovering zero-day vulnerabilities

Evaluating recent Chrome vulnerabilities, Astra discovered two unknown zero-day flaws while achieving an 88% single-pass score on SRE-Bench.

Strict cyber guardrails

To prevent misuse, Astra refuses to build proof-of-concept exploits, while defensive capabilities like patching are gradually opened through OpenAI Daybreak.

Industry-leading safety alignment

In computer-use stress tests, Astra caused just 2.4% misaligned outcomes and recorded a 0% rate of attempting to bypass system safety reviews.

Reduced capability hallucinations

Astra is three times less likely than GPT-5.6 Sol to misrepresent its capabilities or make false claims about what tools it can access.

Real-time misalignment monitoring

OpenAI is deploying production classifiers that monitor Astra's chain-of-thought reasoning and actions, automatically halting unauthorized behavior.

Rollout and API pricing

Astra is rolling out across ChatGPT tiers, API ($10/M input, $50/M output tokens), Azure, and AWS Bedrock, with a 2x Fast Mode available for developers.

At a glance

GPT-6 Astra pairs unprecedented computer use and coding performance with rigorous alignment and real-time safety monitoring.

Read the original on OpenAI

React

Sign in to react and comment.

Comments (0)

Life is short. Keep it sweet. Respect others' opinions and be kind!

    Recommended next

    More decks on ai and related topics.