OpenAI shipped a new flagship model on 3 September 2026, and within a day the internet had produced several hundred summaries of it. This is not one of those. If you already run AI in production — an agent, a CRM sync, an automation layer — the questions that matter are narrower: what does GPT-6 Astra cost per completed task, which of your workloads justify it, and what new failure modes does it introduce into an unattended pipeline. This article answers those, and separates what OpenAI has demonstrated from what OpenAI has asserted.
Quick Answer
GPT-6 Astra is OpenAI’s flagship model, released 3 September 2026. It handles computer use, coding, research, and document creation, with a 1,050,000-token context window and an April 2026 knowledge cutoff. API pricing is $10 per million input tokens and $50 per million output. It is OpenAI’s first model classified at the Critical cybersecurity threshold under its Preparedness Framework.
Last verified: 6 September 2026. We checked every price and rollout claim in this article against OpenAI’s published documentation on that date. Both change frequently.
What is GPT-6 Astra?
Astra is the first model in OpenAI’s GPT-6 generation, announced on Thursday 3 September 2026 and rolled out in stages beginning with a limited set of organisations. It succeeds GPT-5.6 Sol as the flagship. OpenAI has not withdrawn Sol, which remains available in the API at a lower rate.
Unlike the GPT-5.6 generation, which shipped as three tiers — Sol, Terra, and Luna — the GPT-6 line currently consists of Astra and a higher tier called GPT-6 Astra Pro. The two are not the same product: Astra Pro is a separate tier available to users on the Pro, Business, and Enterprise plans, while OpenAI includes standard Astra access within existing subscription allowances.
In short: GPT-6 Astra is a capability step in agentic and computer-use work, sold at a materially higher token price than the model it replaces.
How GPT-6 Astra differs from GPT-5.6 Sol
Three differences matter operationally. First, computer use. OpenAI positions GPT-6 Astra as able to take actions in a browser and on a desktop — filling forms, updating records in business software, running frontend checks — rather than only producing text about those actions.
Second, price. GPT-6 Astra costs $10 per million input tokens and $50 per million output. Sol currently costs $4 and $20 under a promotional rate that OpenAI says runs at least through 21 November 2026, down from a $5/$30 list price. Measured against the rate you would actually pay for Sol today, Astra is 2.5 times more expensive on both input and output. Most launch coverage compared Astra to Sol’s list price and understated the gap.
Third, the safety envelope. OpenAI classifies Astra at the Critical cybersecurity threshold under its own Preparedness Framework, and has deployed production monitoring that can interrupt tasks. That has consequences for automation, covered below.
GPT-6 Astra specifications and pricing
| Attribute | Value |
|---|---|
| API model ID | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | 30 April 2026 |
| Modalities | Text and image in, text out |
| Reasoning effort levels | low, medium, high, xhigh, max |
| Standard input | $10.00 per million tokens |
| Standard output | $50.00 per million tokens |
| Cached input | $1.00 per million tokens |
| Cache writes | $12.50 per million tokens |
| Long-prompt surcharge | Above 272,000 input tokens: 2x input and cache rates, 1.5x output, applied to the whole request |
| Batch and Flex | 50% of Standard rates |
| Fast mode | 2x Standard price for up to 2x Standard speed |
Astra also supports function calling, structured outputs, and MCP, so it slots into an existing tool layer without a rebuild — if you are standing that layer up, our walkthrough of building a production-ready MCP server with Node.js covers the server side. These figures come from OpenAI’s API model documentation. Astra also supports Zero Data Retention for eligible API customers — meaning OpenAI does not store request or response data — which matters if you are handling regulated records.
What the API actually costs per task
Per-million-token rates are hard to reason about. Below are the same rates converted into task costs. All figures assume Standard processing with no caching, and are arithmetic on OpenAI’s published rates rather than measured spend. Re-run them with your own token counts.
| Task | Tokens | Astra | Sol (promo rate) |
|---|---|---|---|
| Contract or report analysis | 40,000 in / 2,000 out | $0.50 | $0.20 |
| Codebase review | 200,000 in / 8,000 out | $2.40 | $0.96 |
| Long-context session | 300,000 in / 10,000 out | $6.75 | $2.70 |
Two levers change these numbers substantially. Batch processing halves them, taking the document analysis to $0.25 — worth using for anything that does not need a synchronous response. Prompt caching, which lets you pay a reduced rate to reuse an unchanged prompt prefix across requests, does more. In a 50-call agent loop with a 30,000-token stable system prompt and 2,000 fresh input tokens per call, caching the prefix cuts the run from roughly $19.75 to roughly $6.63 including the one-off cache write, about two-thirds off.
In other words, on Astra’s rates, prompt caching stops being an optimisation and becomes a design constraint. Structure prompts with a large stable prefix and a small variable suffix, or accept a bill several times higher than necessary.
The long-prompt cliff at 272,000 tokens
This deserves separate attention because it is a discontinuity, not a gradient. Cross 272,000 input tokens and the entire request reprices at twice the input rate and 1.5 times the output rate — not just the tokens above the threshold.
For example, a 270,000-token request with 10,000 output tokens costs about $3.20. Push it to 300,000 input tokens and it costs about $6.75. An 11% increase in input produces a 111% increase in cost.
If you are architecting a long-context agent, this makes retrieval and chunking a cost decision rather than only a latency one, and it makes context budgeting something to enforce in code. We’d suggest a hard ceiling below 272,000 tokens with explicit compaction, and treating any crossing as an alert rather than a silent overage. This is the single most expensive detail in the pricing page and it is easy to miss.
What GPT-6 Astra can do that previous models could not
Computer use and browser control
OpenAI describes computer use as a headline capability for GPT-6 Astra: browser and desktop actions including form completion, CRM record updates, online research, and frontend QA checks. On its own evaluations, OpenAI reports Astra scoring 72.6% on OSWorld 2.0 against 65.7% for GPT-5.6 Sol, and 92.7% on ScreenSpot-Pro against 76.9%. OpenAI also reports that in latency simulations Astra reached the higher OSWorld score in roughly 47% less time per task.
Operationally, therefore, this shifts what is worth automating through a UI rather than an API. Systems without a usable API — legacy portals, government filing systems, vendor dashboards — become candidates for agent-driven work. That said, a UI agent is a fundamentally more brittle integration than an API call, and nothing in this launch changes that. Where an API exists, use it. If you are weighing that trade-off in a CRM context, our guide to building an AI agent for HubSpot CRM covers where structured API access beats screen-level automation, and our comparison of HubSpot MCP versus the HubSpot API for AI agents works through which access method suits which job.
What remains unproven: OSWorld and ScreenSpot are benchmark environments. Neither number establishes reliability against your actual internal tooling, with your actual permission model.
Software engineering and Codex changes
OpenAI reports GPT-6 Astra at 57.9% on Terminal-Bench 4.0 against 37.3% for Sol, and 74.1% on DeepSWE v1.1 against 72.7%. The spread between those two figures is itself informative: the gain is large on some agentic coding tasks and marginal on others.
Alongside the model, OpenAI updated the Codex harness so that Astra can keep notes across context windows rather than repeatedly compacting a long session into a single summary, with earlier context remaining searchable. OpenAI describes this as experimental, and you enable it through Codex configuration. For long refactors and debugging sessions, where compaction loss is a real source of wasted work, this is arguably more useful day-to-day than the benchmark deltas.
Document, spreadsheet, and presentation output
OpenAI states that Astra is trained to follow existing templates and produce documents, slides, and spreadsheets that match a house style, pulling only relevant context into outputs. On its internal AutomationBench evaluation OpenAI reports 41.4% for Astra against 18.1% for Sol.
For teams producing recurring client deliverables from structured data, that is the most directly monetisable capability in the launch. Consistency is the open question. A model that formats correctly nine times in ten still needs a human review step, and the launch materials report no variance figures.
How the benchmark claims should be read
Every headline number above came from OpenAI’s own evaluation environment. That does not make them wrong. It does mean three specific things about how to read them.
First, effort settings. OpenAI states that unless noted otherwise, evaluation scores are the maximum at any reasoning effort. Maximum effort increases latency and token consumption, so a benchmark score and a production cost estimate are not describing the same configuration.
Second, cross-vendor comparisons were produced by OpenAI. The comparison table on OpenAI’s announcement includes competitor models, and OpenAI’s own footnotes disclose the caveats. OpenAI reproduced some competitor results in-house, some results reflect modifications to the evaluation, and for two benchmarks the reported Claude figures come from a variant OpenAI describes as having fewer safeguards. Read the table as one vendor’s account of a competitive field, not as a neutral leaderboard.
Third, some comparisons carry disclosed artifacts. On OpenAI’s internal ExploitBench (June–August 2026) evaluation, OpenAI’s footnote states that Sol’s low score reflects a turn limit real customers would not encounter, and that the same model scored higher when hitting fewer limits. The headline gap between models on that benchmark is partly an artifact of the harness, and OpenAI says so.
Still, none of this is a debunking. OpenAI documented every one of these caveats itself, in public, in footnotes. The point is that the footnotes materially change several readings, and almost no coverage reproduces them.
Cybersecurity, alignment, and the safety trade-off
OpenAI classifies GPT-6 Astra as reaching the Critical cybersecurity threshold under its Preparedness Framework. This is OpenAI’s internal risk designation applied by OpenAI, not an external certification. In practice it means the publicly available model refuses advanced offensive tasks such as writing proof-of-concept exploits, while OpenAI plans to extend less restrictive safeguards to vetted organisations through its Daybreak programme for defensive workflows including malware analysis and detection engineering. Full detail is in the GPT-6 Astra system card.
Two disclosures deserve attention from anyone running unattended automation.
Why a GPT-6 Astra API task can stop outright
The first disclosure is an operational fact, not a risk assessment. OpenAI has deployed misalignment monitoring in production for Astra-class models, and states plainly that these extra safety checks can sometimes slow, pause, or stop legitimate work. In ChatGPT or Codex, a paused task may prompt you to review before continuing. In the API, the task stops. That is a new failure mode. If you run a scheduled pipeline where a step can now terminate for reasons unrelated to your code, your error handling needs to distinguish a safety stop from a timeout or a malformed response, and your retry logic needs to not loop forever against it. Teams building unattended AI agents and automation should treat this as a first-class case in the pipeline design, not an edge case discovered in production. If you run scheduled jobs, the pattern we describe for an n8n agent that updates HubSpot is the right place to add that branch, because the write step is where a silent stop does the most damage.
What OpenAI disclosed about monitorability
The second disclosure cuts against OpenAI’s own launch narrative. In its safety overview, OpenAI reports that Astra’s written reasoning was harder to monitor than Sol’s under tests that explicitly instructed the model to evade monitoring. OpenAI attributes this to Astra exercising greater control over its written reasoning and solving problems in fewer written steps, notes that its broader alignment evaluations show Astra less likely than Sol to violate safety restrictions overall, and says it takes the trend seriously. A vendor publishing a result that complicates its own release is worth crediting, and worth reading carefully rather than either dismissing or amplifying.
On the other side of the ledger, OpenAI reports Astra as significantly more robust to prompt injection than Sol — the attack where instructions hidden in retrieved content hijack an agent’s behaviour. For any agent that reads untrusted web pages or inbound email, that is the more relevant safety property, and an improvement there is worth more than most of the benchmark deltas.
Who can access GPT-6 Astra, and when
Access rolled out in stages rather than all at once, which is why availability varies by account. OpenAI began with a limited set of organisations, then expanded to ChatGPT Plus, Pro, Business, and Enterprise users, the OpenAI API, Microsoft Azure, and Amazon Bedrock. Astra is generally available in Microsoft Foundry.
If Astra is missing from your model selector, check three things before assuming a fault. On Enterprise plans, access is off by default and a workspace administrator has to enable it. Astra Pro is a separate tier limited to Pro, Business, and Enterprise. And staged rollouts do not reach every account simultaneously. OpenAI includes Astra usage within existing subscription allowances, and sells additional usage as credits.
Should you migrate? A decision framework
The useful question is not whether GPT-6 Astra is better. It is which of your workloads change outcome at 2.5 times the token cost. Four common cases:
- High-volume classification, extraction, and field mapping. Do not migrate. At 10,000 records a month with 1,500 input and 300 output tokens each, Astra costs roughly $300 while GPT-5.6 Luna costs roughly $6.60 — about 45 times more for work where the cheaper model is already at ceiling accuracy. Luna and Terra remain available and were repriced downward on 30 July 2026.
- Templated content and record enrichment. Do not migrate wholesale. Route by ambiguity instead: send the clean records to a cheap tier and escalate only the ones that fail a confidence or validation check. Most workflow automation spend sits here, and this is where a routing layer pays for itself fastest. Our guide to connecting ChatGPT to HubSpot CRM using MCP shows the plumbing that makes swapping the model behind a workflow a configuration change rather than a rewrite.
- Multi-step agentic tasks with irreversible actions. Migrate and test. Creating deals, sending invoices, updating production records — anywhere a wrong answer costs more than the model does. The deciding variable is the cost of a mistake, not the difficulty of the task.
- Anything that currently requires a human because the UI has no API. Test first. This is where computer use genuinely opens new ground, and also where reliability is least established.
The deciding variable is cost of error, not task difficulty
The general heuristic: Astra earns its price where inputs are ambiguous, steps are many, and errors are expensive. It does not earn its price on volume. If you are running one model for everything, the highest-return change available right now is a routing layer, not a model upgrade — and that holds whether or not you adopt Astra at all. The same logic applies to CRM integration work, where the majority of operations are deterministic and do not need a reasoning model at any tier.
Limitations and open questions
Several things are not yet known, and it is worth being explicit about them.
- No independent evaluation has confirmed the headline results. Every figure in the launch materials is OpenAI’s, run in OpenAI’s environment, at maximum effort unless noted.
- OpenAI reports token-efficiency gains on several evaluations, but has not published enough data to establish whether those savings offset a 2.5x rate increase on real workloads.
- Nobody has yet measured production reliability over long agentic sessions outside benchmark harnesses.
- OpenAI has not published how often safety checks interrupt normal API use, and says it is still iterating to reduce unnecessary interruptions.
- Rollout status changes daily, and the Sol promotional rate that anchors every price comparison here expires.
- Better reasoning does not eliminate fabrication. If your system needs factual reliability, grounding and verification still do that work — the failure modes covered in our guide to preventing AI chatbots from hallucinating are unchanged by a model upgrade.
Frequently Asked Questions
Is GPT-6 Astra worth the price increase over GPT-5.6 Sol?
It depends entirely on workload. At 2.5 times Sol’s current rate, Astra pays off where an error is expensive or a task previously needed a human. For high-volume, low-ambiguity work it does not pay off, and OpenAI has not published data showing its token-efficiency gains offset the rate increase.
Why can’t I see GPT-6 Astra in my ChatGPT account?
Access rolled out in stages rather than all at once. On Enterprise plans, Astra is off by default and a workspace administrator must enable it. If it is missing from your model selector, check your workspace policy and plan tier before assuming an account problem.
What is the difference between GPT-6 Astra and GPT-6 Astra Pro?
Astra Pro is a separate, higher tier available to users on Pro, Business, and Enterprise plans. OpenAI includes standard Astra access within existing subscription allowances and sells additional usage as credits. They are distinct products, not settings on the same one.
What does the Critical cybersecurity classification mean in practice?
It is OpenAI’s internal risk designation under its own Preparedness Framework, not an external certification. Practically, the publicly available model refuses advanced offensive tasks such as writing proof-of-concept exploits, and OpenAI instead routes less restrictive access to vetted organisations through its Daybreak programme.
Can GPT-6 Astra control a computer?
Yes. OpenAI positions computer use as a core capability covering browser and desktop actions — filling forms, updating records in business software, running checks on web applications — and reports state-of-the-art results on its own computer-use evaluations. Reliability against specific internal tooling remains untested publicly.
How does prompt caching change GPT-6 Astra’s cost?
Substantially. Cached input bills at $1 per million tokens against $10 standard, with cache writes at $12.50. On a repeated agent loop with a large stable prompt prefix, caching can cut a run by roughly two thirds. At Astra’s rates, prompt structure is a cost decision.
Does GPT-6 Astra support Zero Data Retention?
OpenAI states that Astra supports Zero Data Retention for eligible API customers, meaning request and response data is not stored. OpenAI determines eligibility rather than letting you select it in the API, so confirm your account’s status directly before assuming it applies to regulated workloads.
Can an API task fail because of Astra’s safety monitoring?
Yes. OpenAI has deployed misalignment monitoring in production and states that safety checks can slow, pause, or stop legitimate work. In ChatGPT and Codex you may be asked to review; in the API the task stops. Unattended pipelines need to handle this explicitly.
What to do next
GPT-6 Astra is a real capability step, priced accordingly, and the interesting question is workload selection rather than wholesale adoption. The genuinely new thing for most teams is delegating multi-step browser and desktop work that previously needed a person. The thing not yet proven is how the reported gains hold up outside OpenAI’s evaluation environment.
So produce evidence instead of an opinion. Pick one existing automation — one you already run and already have quality data on. Run it on Astra and on your current model with identical inputs. Compare accuracy against cost per completed task, not cost per token. That takes an afternoon and gives you a decision.
If it would help to have someone look at where model routing would cut cost across your current stack, we’re happy to review it.





