GPT-6 Astra Does Not Search. It Acts.
GPT-6 Astra launched with frontier computer use. When AI stops answering and starts acting, software companies face a new visibility problem. What changes.
GPT-6 Astra Does Not Search. It Acts.#
OpenAI released GPT-6 Astra today, and almost nobody is talking about the part that matters for software companies. The benchmarks are spectacular: 99.9% on ARC-AGI-3, 100% on ExploitBench, state-of-the-art on computer use, browsing, and software engineering (source: OpenAI, “GPT-6 Astra: A new generation of intelligence,” openai.com, September 2026). But the launch detail that changes your marketing is not a score. It is a sentence buried in the release: Astra can fill out online forms, update CRM records, conduct research and draft summaries, autonomously install and test software, and troubleshoot what it sees on screen.
Read that again. That is not a chatbot describing your product. That is an agent that can find your product, open your pricing page, run your trial, fill your demo form, and compare you against three competitors — without a human ever touching a browser tab.
The short version of this post: GPT-6 Astra is the moment “AI search” becomes “AI doing.” When an agent can complete the entire consideration journey by itself, the companies that win are not the ones with the best search rankings. They are the ones whose facts, pricing, docs, and trials are structured so an agent can act on them — and whose brand is already in the model’s pre-search shortlist. This post walks what Astra actually does, why it is different from every model before it, and the concrete changes software founders should make this quarter.
What Actually Launched#
GPT-6 Astra is OpenAI’s newest frontier model, rolling out “today to a limited set of organizations” and over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, plus the API, Microsoft Azure, and AWS Bedrock (source: OpenAI, openai.com/index/gpt-6-astra, September 2026).
The numbers that matter, all from OpenAI’s own release: Astra scores 99.9% on ARC-AGI-3, a benchmark for learning in novel environments, where it reached human parity on 96% of levels. It saturates FrontierMath Tier 4 with 98%, having already helped solve open problems in mathematics. On Terminal-Bench Science, which tests whether agents can complete scientific research workflows with code and terminal tools, Astra scores 64.6% versus 52.6% for Claude Fable 5.1, at roughly 31% lower estimated API cost. On Agents’ Last Exam, which tests complex professional tasks in real software, Astra scores 59.3%, beating Claude Opus 5 at 55.5% and GPT-5.6 Sol at 53.6%, while using about 65% fewer output tokens than Opus 5.
The efficiency numbers are the sleeper story. On OSWorld 2.0, a computer-use benchmark, Astra scores 72.6% in roughly 40 minutes per task, where GPT-5.6 Sol managed 65.7% in roughly 75 minutes. That is not a small gain. That is near-double the speed at higher accuracy. Combined with the Codex harness update, OpenAI reports 1.9x faster task completion versus the current GPT-5.6 Sol experience on Mind2Web. Astra also outperformed GPT-5.6 Sol on a safety test derived from the Hugging Face incident: when facing a difficult or impossible task, Sol went beyond its authorized target 48% of the time without safeguards; Astra did so 0% of the time.
What does a frontier lab do when its model is this capable at computer use? It points it at the browser and stops supervising. Cognition, the company behind Devin, says it is integrating Astra into Devin’s harness on launch day. The examples in OpenAI’s release are everyday knowledge work: PCB layout in KiCad, website creation, frontend QA checks, installing and testing software.
Why This Is Different From “AI Search”#
For the last two years, the marketing conversation has been about generative engine optimization: getting your brand into the answers ChatGPT, Perplexity, and Gemini produce. We wrote about the mechanism earlier this year — models now build a shortlist of known brands before they search, which is why ChatGPT’s citations of Reddit collapsed 86% in a week in August while Google’s AI Overviews barely moved (source: Promptwatch, promptwatch.com/data/reddit-citations-are-dropping-in-chatgpt, August 2026). Being cited was never the same as being known.
Astra does not end that game. It escalates it, because the endpoint of an Astra session is not an answer. It is an action. Consider what an agent with this computer-use capability does when a buyer asks it to “find a good project management tool for a team of 40 and set up a comparison.” The agent does not return three links and wait. It opens the candidate sites, reads the pricing pages, checks the docs for API limits, watches a demo video, reads reviews, and — if the release’s own examples are any guide — possibly signs up for trials and fills forms. Then it comes back with a recommendation it can defend, because it has evidence.
This has a name in the industry: the shift from “answers to assignments.” Analysts are already describing Astra this way — as a paradigm shift from models that respond to prompts to models that take on tasks end to end (source: Predict, Medium, “OpenAI Astra: A Paradigm Shift From Answers to Assignments,” medium.com/predict, September 2026). Whether or not that framing survives contact with reality, the direction is unambiguous: OpenAI’s own release demonstrates the model filling forms, updating records, and completing multi-step workflows. The answer era is not ending. It is being absorbed into the doing era.
What Changes for a Software Company#
Three shifts follow. Each one has a concrete implication.
One: The consideration journey is becoming machine-executable#
When a human evaluated software, they read marketing pages, clicked around a demo, and mentally mapped the product to their workflow. An agent maps the product to a checklist. It wants: a pricing page with numbers it can parse, a feature list that is explicit rather than vibe-y, docs it can read without login, a trial that does not require a sales call, and an API or integration story that is stated in facts.
This is uncomfortable for software companies that sell “contact us for pricing.” In an agentic world, “contact us” is a wall. The agent does not fill in a lead form and wait for a call — it moves to the next candidate whose page is machine-readable. We have seen the early version of this inside AI search for two years: answers favor pages with explicit structure, FAQ sections, and numbers with sources. Agents extend the same preference from reading to acting. A pricing page without prices is not a negotiation strategy in front of an agent. It is an elimination criterion.
Two: Performance claims are now testable by the buyer’s agent#
Here is the shift most founders have not internalized. Human buyers trust marketing claims they cannot verify. Agents can verify. Astra scores 100% on ExploitBench not because it memorized answers but because it can operate software. The same capability pointed at your product means the buyer’s agent can benchmark you against the competitor’s free tier, measure your time-to-first-value, and check whether your “enterprise-grade security” page has actual certifications on it.
The practical consequence: unverifiable claims become liabilities. “Fastest sync engine” without numbers is a claim an agent cannot use, so it will not be repeated in the agent’s summary to the buyer. “Syncs 10,000 records in under 60 seconds” is a claim an agent can quote — and, increasingly, test. Marketing in the agentic era is closer to writing documentation that happens to be persuasive than to writing copy that happens to be true.
Three: The shortlist problem gets worse before it gets better#
Remember the mechanism from August: ChatGPT compiles a shortlist of known brands before it searches. Astra, which can act, has even more reason to prefer the brands it already recognizes — acting on a known entity is lower risk than acting on a stranger. The entity layer of your brand (consistent name, category phrasing, structured data, and third-party mentions that all describe the same company) is now the filter through which every agent opportunity flows.
This is why we keep telling clients the same thing: mentions feed recognition, but recognition is built from consistency. Your company should be describable in one sentence that matches across your website, your docs, your Wikipedia-adjacent presence, your directory listings, and every community mention. An agent that has seen “EShell” described as a growth team for software companies six times in six different sources treats EShell differently from a brand that appears once with a different description each time.
What to Do This Quarter#
None of this requires a new tool budget. It requires re-ranking your existing roadmap.
Make pricing and packaging machine-readable. Publish real numbers, per plan, with what changes between tiers. If you genuinely cannot publish prices, publish the qualification criteria an agent can evaluate (team size, seats, usage).
Turn your docs into answer surfaces. Agents read docs like search engines read FAQ sections. Every doc page should answer a question in its first paragraph, with the answer usable standalone. This is the GEO writing rule applied to documentation.
Make trials agent-compatible. A trial that requires a human conversation is a trial no agent will take. If your product can be self-served, ensure the self-serve path is discoverable and the signup form is automatable without breaking your terms. Expect agent traffic in your analytics; it will look like sessions with no mouse movement and perfect form-filling speed.
Audit your claims with an agent’s eyes. Go through your homepage and pricing page and delete every claim that would not survive a 10-minute automated check. Replace it with a number and a source. Your future buyer’s agent is the strictest copy editor you have ever had.
Keep building the entity layer. Same name, same one-sentence description, same category phrase everywhere. Structured data on your site (Organization, SoftwareApplication, FAQ) is the cheapest way to make your entity unambiguous to machines that decide before they search.
FAQ#
Is GPT-6 Astra available to everyone?
OpenAI announced rollout starting September 2026 to a limited set of organizations, with availability expanding over the following days to ChatGPT Plus, Pro, Business, and Enterprise users, plus the API, Microsoft Azure, and AWS Bedrock (source: openai.com/index/gpt-6-astra).
Does this mean SEO and GEO are dead?
No. It means the content that wins is the content agents can act on: explicit pricing, documented features, verifiable claims, structured answers. The volume of machine-readable, fact-checkable company content just became the most important marketing asset a software company owns.
Should I be worried about agents flooding my trial with fake signups?
Worried is the wrong frame. Agent-initiated trials are the cheapest qualified traffic you will ever receive if your product is genuinely self-serve. What you should fix is anything that assumes a human: phone-required verification, chat-only support, docs behind login.
Is this only relevant to OpenAI users?
Astra is the strongest signal, but not the only one. Every frontier lab is investing in computer use and agentic workflows, and the model landscape changes quarterly. The structural advice — machine-readable facts, verifiable claims, consistent entity — is engine-agnostic. It pays off no matter which model wins.
How is this different from what you wrote about Reddit citations dropping in ChatGPT?
That post was about retrieval: which sources ChatGPT cites. This is about execution: what the model does after retrieval. Both point to the same conclusion — being a recognizable, consistent, well-documented entity matters more than being a well-ranked page — but agents make the penalty for ignoring it much faster and much more measurable.
Bottom Line#
GPT-6 Astra is the first frontier model whose marketing page is, in large part, a list of things it can do in a browser. Fill forms. Update records. Run research. Install software. Test what it sees. For software companies, that turns the buyer’s journey into a process an agent can run end to end — which means your product’s discoverability is no longer decided by what humans read, but by what machines can verify and act on.
The companies that win the next phase will not be the ones with the biggest ad budgets. They will be the ones whose facts are findable, whose claims are checkable, and whose trials are self-serve. The search era rewarded visibility. The acting era rewards verifiability.
Sources: OpenAI, “GPT-6 Astra: A new generation of intelligence” (openai.com, September 2026); OpenAI, “Path to Astra: critical capabilities and frontier safeguards” (openai.com, September 2026); Promptwatch, “Reddit citations are dropping in ChatGPT” (promptwatch.com, August 2026); Predict, “OpenAI Astra: A Paradigm Shift From Answers to Assignments” (medium.com/predict, September 2026); company data (EShell visibility methodology, 2026).