If you have been paying for ChatGPT Plus and quietly wondering whether it earns its place, the answer changed this week. Not because the subscription changed, but because of what is now sitting inside it.

OpenAI released GPT-6 Astra, and the interesting part for a small business is not the benchmark scores everyone is screenshotting. It is where the model lives. Astra is included in the existing paid plan allowances rather than being something you top up separately. That one detail does more to the value of a Plus subscription than any accuracy percentage.

I went through the whole launch page looking for the parts that actually touch a working business, ignored most of it, and then checked what it costs to run against Claude. Here is what I found.

Most of an AI launch does not apply to you

The Astra announcement opens with about two and a half minutes of talking about how the model gets used. Very little of it matters if you run a small business.

That is not a criticism of OpenAI. It is how these launches work. They are written for researchers, developers, enterprise buyers and the general press all at once, so the majority of any launch page is aimed at somebody who is not you. Prime number gaps, reverse engineering software binaries, printed circuit board layout in KiCad. Genuinely impressive, completely irrelevant to a business owner trying to get their marketing done.

The trick is knowing which handful of numbers to look at and ignoring the rest. For a small business, there are really only a few questions worth asking of any new model:

  • Can it finish a multistep job rather than just describing one?
  • Can it hold a long professional task together from start to finish?
  • Can it use a computer and work across the apps I already use?
  • Can it do research, competitor work and lead work?
  • Can it produce documents and decks that match how I already work?
  • What does it cost me to actually run?

Everything else is interesting reading. Those six decide whether a subscription is worth paying for.

The AutomationBench chart on OpenAI's GPT-6 Astra launch page, plotting accuracy against API cost.
OpenAI's own AutomationBench chart, plotting accuracy against API cost. The vertical axis is what most people screenshot. The horizontal axis is the one that decides whether you can afford to use it.

The one benchmark that actually maps to running a business

AutomationBench is the number I care about most, and it is the one I would point any business owner at first.

It tests whether an agent can complete multistep business workflows across applications. Not answer a question, not summarise a document, but finish a job that has several stages and several places it could go wrong. That is exactly the work most small businesses want to hand over, and it is exactly where AI tends to fall apart in practice.

On OpenAI's published figures, GPT-6 Astra completes 41.4% of those tasks. Claude Fable 5.1 completes 31.4%. Claude Opus 5 completes 26.9%.

Comparison of AutomationBench accuracy: GPT-6 Astra 41.4%, Claude Fable 5.1 31.4%, Claude Opus 5 26.9%.
A ten point gap on the benchmark that most closely resembles real business work. Worth noting that none of these models finish even half the tasks.

Ten points is a real gap. But read the number properly before you get excited: the best model available finishes fewer than half of these jobs. Anyone telling you that AI now runs a business unattended is selling something. What the number actually says is that automation is getting meaningfully more reliable, and that the gap between the leading models on this specific kind of work is no longer small.

The other business-facing scores follow a similar pattern. On Agents' Last Exam, which tests long professional tasks carried start to finish, Astra reaches 59.3% against Claude Opus 5 at 55.5%. On internal design tasks, the kind that produce documents and presentations, Astra hits 50.0% against Claude Fable 5's 35.8%. On internal data science tasks it reaches 40.9% against 34.7%.

On research work the picture is different, and this is worth saying plainly because it cuts against the headline. BrowseComp measures the sort of work I do constantly, competitor analysis and lead research. Astra scores 91.5%. Claude Opus 5 scores 90.8%. That is not a gap, that is a rounding difference. If research is the main thing you use AI for, this launch does not give you a reason to move.

Screen reading is the quiet upgrade

One score jumped more than any other and it got almost no attention.

ScreenSpot-Pro measures whether a model can actually read what is on a screen and work out where things are. It sounds dry. It is the foundation for every "AI uses your computer on your behalf" promise, because a model that cannot reliably tell a button from a label cannot be trusted to click anything.

Astra scores 92.7% on it. The previous model, GPT-5.6 Sol, managed 76.9%.

That is a much bigger jump than any of the headline numbers, and it is the one that decides whether agents doing real work on your machine move from a demo to something you would actually leave running. If you have watched an AI agent confidently click the wrong thing, this is the number that was holding it back.

There is a related figure worth knowing. On OSWorld, which tests computer use directly, Astra reaches 72.6% in roughly 40 minutes per task, against the previous model's 65.7% in roughly 75 minutes. Better result, and about 47% less time to get there. Speed matters more than it sounds when you are paying per task.

What it costs, which is the part that decides it

Here is where a Plus subscription starts looking different.

Artificial Analysis publishes a cost per Intelligence Index task, which is a weighted average of what it costs to run one standardised task, broken down by token type. Lower is better. It is a reasonable proxy for what a model costs you in practice rather than what its list price says.

The Artificial Analysis cost per Intelligence Index task chart, showing model costs side by side.
Artificial Analysis, read on release day. Cost charts move, so treat this as a snapshot rather than a fixed price.

On the day Astra launched, GPT-6 Astra at max effort worked out at $1.67 a task. Claude Fable 5.1 at max effort worked out at $3.69.

Cost per Intelligence Index task: GPT-6 Astra max at $1.67, Claude Fable 5.1 max at $3.69.
Less than half the cost per task, for better scores on most of the business-facing benchmarks. This is the comparison that matters more than any single accuracy figure.

Claude is the most expensive of the lot for most things, and it has been for a while. That has been a fair trade when Claude was clearly ahead. It is a harder trade when it is behind on automation, design and data work, and costs more than twice as much per task.

For completeness, OpenAI's published API pricing for Astra is $10 per million input tokens and $50 per million output tokens, with separate rates for cache reads and writes. If you are building on the API rather than working inside the app, those are the numbers to plan against.

There is one outside voice on the cost question worth quoting, because it comes from a company running these models at volume rather than from a benchmark. Higgsfield, which generates video and images at scale, tested Astra before launch. Their co-founder Alex Mashrabov said it executes their most complex creative workflows "while using up to 20% fewer tokens than other models we've tested."

That is a different measurement from the per-task figure, and it points the same way. A model that needs fewer tokens to finish the same job costs less to run even when the headline rate per token is identical. For anyone paying per use rather than working inside a flat subscription, efficiency compounds quietly in a way that list pricing never shows you. It is also the sort of claim a paying customer has no incentive to exaggerate downward.

The structural difference nobody is talking about

The per-task cost is the headline. The plan structure is the thing that actually changes the decision.

Astra usage is included within existing subscription allowances. Users and businesses can buy credits for additional usage on top, but the baseline comes with the plan you are already paying for. Claude, by contrast, keeps pushing you toward buying more token usage or going through API access when you run past what you have.

That is the big drawback Claude has here, and it is the reason I think people start swinging back toward ChatGPT. Not because of a benchmark. Because of how the bill arrives.

There is a real difference between a tool that costs a predictable amount every month and a tool that costs a predictable amount every month plus whatever you happened to use. For a small business trying to forecast anything, the first one is far easier to live with, even at a similar headline price. Most owners I speak to would rather have a slightly worse tool with a bill they can predict than a slightly better one that surprises them.

Which plans actually get it

This is where the "is ChatGPT Plus worth it" question gets its clearest answer.

GPT-6 Astra is rolling out to Plus, Pro, Business and Enterprise users, along with the OpenAI API, Microsoft Azure and Amazon Bedrock. Free accounts are not on that list. Pro, Business and Enterprise plans additionally get access to GPT-6 Astra Pro, and enterprise administrators have to switch it on for their workspace because it is off by default at launch.

So the honest answer to whether Plus is worth paying for is that the same subscription now carries a materially better model for business work than it did last week, at no change in what you pay. That is not a small thing. It is the first time in a while that the value of the paid tier moved without the price moving.

If you are on the free tier and you have been waiting for a reason to upgrade, this is a more concrete one than most. You are not paying for a slightly faster answer any more. You are paying for the model that finishes multistep jobs at 41.4% instead of the one that does not have access to them at all.

What I am actually doing about it

I do not have access to Astra yet. It rolled out to a limited set of organisations first and works its way out over the following days, which is normal.

That has not stopped me planning around it. Everything in my setup is built so I can move between AI tools without rebuilding the whole thing, which is exactly the situation this kind of release is for. If the numbers hold up when I get hands on it, moving more of my system across to ChatGPT is a genuine option, and the deciding factor will be the cost structure rather than the benchmark scores.

A word of caution before you move anything. These are OpenAI's own published figures on OpenAI's own launch page, measured at maximum effort settings. Vendor benchmarks are not neutral, effort settings change the numbers substantially, and the low effort results are far less impressive than the maximum effort ones. On design tasks the difference between the low and high settings on that chart is the difference between something unusable and something good.

Worth knowing too: on the Artificial Analysis Intelligence Index, which OpenAI reprints on its own page, Claude Fable 5.1 still scores higher overall than Astra, 65.7 against 61.2. Astra's wins are specific to business-shaped tasks rather than across the board. That is a narrower claim than most of the coverage is making, and it happens to be the claim that matters if you run a business rather than a research lab.

Frequently asked questions

Is ChatGPT Plus worth it for a small business? More so than it was last week. GPT-6 Astra is included in the Plus allowance, and on OpenAI's published figures it completes 41.4% of multistep business workflow tasks against Claude Fable 5.1's 31.4%. The subscription price has not changed, so what you get for it has gone up.

Does the free ChatGPT plan get GPT-6 Astra? No. OpenAI lists Plus, Pro, Business and Enterprise. Free accounts are not included. Pro, Business and Enterprise also get GPT-6 Astra Pro.

Is ChatGPT cheaper than Claude to run? On the day of release, yes, by a good margin. Artificial Analysis put GPT-6 Astra at max effort at $1.67 per Intelligence Index task and Claude Fable 5.1 at max effort at $3.69. Cost charts move, so check before you commit.

Should I cancel Claude and switch everything to ChatGPT? Not on one launch. Research scores are almost level between the two, Claude Fable 5.1 still leads on the overall Intelligence Index, and these are vendor figures at maximum effort. The sensible move is to keep your setup portable so you can move when the evidence is yours rather than the vendor's.

What does GPT-6 Astra cost on the API? OpenAI's published standard pricing is $10 per million input tokens and $50 per million output tokens, with separate rates for cache reads and writes. A faster processing mode is available at roughly twice the speed for twice the price.


Disclosure: Some links in this article are referral links. If you use one, I may earn a commission at no extra cost to you. I only link to tools I actually use.

How this was made: The facts, numbers and examples in this article come from my own work and videos. I use AI to help draft and structure the writing, then I review and edit every line myself before it goes live.

Wayne St Ledger

About the author

I'm Wayne St Ledger, and I run St Ledger Marketing. I started out in construction as a plasterer, then went on to a Bachelor's in Marketing and a Master's in International Business, which is where my bias toward results over talk comes from. I help business owners build marketing systems they actually understand, using AI in grounded, practical ways rather than hype. I also run a YouTube channel on marketing and AI, and The AI Marketing Hub, my own community for business owners and marketers putting AI to work properly.