Sep 4, 2026

Tools

GPT-6 Astra: What OpenAI's New Model Means for Business

A spiral of small beads arranged on paper

What changes when AI can use the software you already run

On 3 September 2026, OpenAI released GPT-6 Astra. The headline capability is not that it answers better. It is that it can now operate the software a business already runs.

Every AI tool until now has sat next to the work

For three years, using AI has meant having a conversation. A person types a question, the AI writes back, and then the person opens the right piece of software and does the job themselves.

That shape held for everything built on top of it. The AI drafts the quote, somebody puts it into the system. It summarises the call, somebody updates the record. It writes the email, somebody sends it.

The AI never touched the business. It sat beside it and offered opinions.

This is why a lot of businesses tried AI, found it genuinely impressive, and could not say afterwards what it had saved them. Something that produces good advice still leaves every piece of the actual work where it was. The bottleneck was rarely knowing what to write. The bottleneck was the forty minutes of clicking that followed.

What changed on 3 September 2026

OpenAI released GPT-6 Astra to a limited set of organisations, with wider release to paid ChatGPT plans and the developer interface over the following days.

The company framed the launch around one capability above the others. The model operates a computer. Not through special connections built for it in advance, but by looking at the screen, reading what is there, moving the pointer, clicking, typing and moving between one program and the next.

Greg Brockman, OpenAI’s president, described it as being able to zip through spreadsheets, fill out forms and navigate across web pages.

This is the part that matters, and it is easy to lose underneath the louder claims about intelligence and the argument over whether any of this counts as artificial general intelligence. That argument will run for a year. The capability is available now.

Astra is also not alone. Anthropic and Google have been building the same thing. The date matters less than the direction, which every major provider is now moving in at once.

Nine tasks in ten, and roughly half the time

There is a standard test for this called OSWorld. It hands a model a real computer and a list of ordinary jobs, then counts how many get finished.

Astra finished 72.6% of them, against 65.7% for the model it replaces. Average time per task fell from 75 minutes to 40.

One caveat belongs with that figure. It is OpenAI’s own number, published by OpenAI, on the day it launched a product, and the tests were run at maximum effort, which flatters both the score and the timing. Nine tasks in ten is also not ten in ten.

A useful way to picture the change: earlier versions behaved like a capable temp who has never seen your systems and keeps clicking the wrong thing. This one gets it right most of the time, and gets there faster.

Getting information into the system of record

The tasks OpenAI demonstrated were ordinary office work. Updating customer records, filling in forms, organising a calendar, building spreadsheets, preparing documents, creating a listing on eBay.

The first job worth handing over is the one nearly every business has. A call happens, an email arrives, an enquiry form comes in, and somebody has to type it into the CRM.

What makes this suit the new capability is that no integration project is required. The model types into the same boxes a person types into. A business that has put off connecting its systems because the integration was quoted at several thousand pounds now has a second route to the same outcome.

Moving information between systems that do not talk

A business running three or four tools that were never designed to work together has somebody copying between them by hand.

Reading from one screen and typing into another is precisely what this capability is. It does not care whether the two systems have an interface for talking to each other, because it is not using one.

This is the least glamorous work in most companies and some of the most expensive, because it scales directly with volume. Every new customer adds another round of it.

The weekly grind nobody wants

Pulling the same report every Monday morning. Checking a supplier portal for updates. Reconciling two lists that should match and occasionally do not.

Work that is trivial each time it happens and expensive across a year.

The pattern connecting all three of these jobs is that somebody already knows exactly what to do and simply has to do it. That is the work this suits. Judgement calls, negotiation and anything where being wrong is costly all stay where they are.

It costs more, and it is not more intelligent

Three things deserve stating plainly before anyone budgets for this.

Businesses on a paid ChatGPT plan get access within their existing allowance, with additional usage bought as credits. Built directly into a system, it is priced at two and a half times the model it replaces.

More importantly, it is not more intelligent. Artificial Analysis, an independent benchmarking firm, scored Astra at 61 on its intelligence index, level with the model it replaces. The gain is specifically in doing things, not in being cleverer about them. A business hoping for better answers to hard questions is paying two and a half times more for the same quality of answer.

It also still gets things wrong. OpenAI’s own measure of factual errors improved from 12.2% to 4.2%. Better, and not zero. One in twenty-five is fine for a first draft and not fine for anything that sends, files or pays without a person seeing it.

Three things have to be true before handing work over

For three years the question was whether AI could produce something good enough to use. For most everyday work, that is settled. The question now is which jobs are worth handing to something that acts on its own.

Three things have to be true.

The job has a right answer. Not a judgement call, not a negotiation, not something where two reasonable people would do it differently.

Being wrong is cheap, or gets caught quickly. A mistake that surfaces immediately is a nuisance. A mistake that surfaces in three months is a different category of problem.

A person sees the result before a customer does. Anything that sends money, makes a promise or creates a record a regulator could ask about needs somebody looking at it.

Anything meeting all three is worth looking at now. Anything failing one should wait, and that is not a temporary position while the technology improves. It is what handing work to anybody, human or otherwise, has always required.

The useful question before automating a task is no longer whether AI is clever enough to do it. It is whether the job has a right answer that somebody is currently clicking through by hand.

Sources: OpenAI GPT-6 Astra announcement and benchmark tables, 3 September 2026. Independent benchmarking by Artificial Analysis. Reporting by TechCrunch, VentureBeat and Fortune. Benchmark analysis by Vellum and DataCamp. Every performance figure except the intelligence index originates with OpenAI.

FREE EMAIL BRIEFING

The Tustra Briefing

AI in plain English for UK business.

One considered briefing covering what changed, why it matters, how businesses are using it, what to be cautious about and one practical action worth considering.

Important developments without daily noise

Practical UK business context

Honest case studies, security and adoption guidance

Free to join. Unsubscribe at any time. We will not sell your information.