Fable 5.1 vs GPT-6 Astra: What I Tell a Boardroom
16 Sep 2026 · 5 min read
Anthropic shipped Claude Fable 5.1 on September 1. OpenAI shipped GPT-6 Astra on September 3, with general availability the day after. Within a week, three different CEOs sent me the same question in three different ways: which one should we use?
It's the right question, asked in the wrong shape. So here's the shape I give it back in.
The list price is a coincidence. The bill is not.
Both models cost $10 per million input tokens and $50 per million output tokens. That's not an accident of the market; that's two companies pricing against each other. If you stop reading there, you conclude they cost the same.
They don't. The difference is in what happens when a model reads something it has already read.
Almost every serious business use of these tools is repetitive by design. The same company knowledge base, the same brand guidelines, the same contract template, read again and again with a small new question at the end. That repeated part is served from cache, and cache reads are where the two diverge: $0.25 per million on Fable 5.1, $1.00 on Astra. On a request that runs past 272,000 tokens of context, OpenAI also adds a surcharge. Anthropic doesn't.
Then it flips. Independent evaluator Artificial Analysis measures what a model actually spends to complete a task, not what it charges per token, and on their index Astra came out at roughly $1.67 per task against $3.76 for Fable 5.1. Astra gets to its answer with fewer words.
So: a loop that hammers the same large context is cheaper on Fable 5.1. A one-shot question that needs a sharp answer is cheaper on Astra. Which one you have is a question about your workflow, not about the models.
The benchmarks disagree, on purpose
OpenAI's own table has Astra ahead almost everywhere. Math, science, coding, computer use, security. Some of the gaps are wide: 97.6% against 87.8% on FrontierMath Tier 4, 41.4% against 31.4% on AutomationBench.
Artificial Analysis, who don't work for either company, have Fable 5.1 first of 202 models on their Intelligence Index at 66, with Astra eighth at 61. Their coding agent index also leans Fable, 70 against 67. And on Humanity's Last Exam with tools, Anthropic's number, 65%, beats OpenAI's 57.2%.
I don't think either side is lying. I think each side chose the tests that flatter its model, which is what marketing departments are for. When I put both tables on a screen, the room usually goes quiet, and then someone says: so it depends. Yes. It depends. That's the honest answer, and it was true before this week.
What the tables agree on is more useful than where they differ. Astra is genuinely strong at driving a computer: clicking through interfaces, filling forms, operating software that has no API. Fable 5.1 is genuinely strong at long, deep reasoning across a lot of material. If your problem is "make the machine do the boring thing in our old ERP," look at Astra. If your problem is "read four hundred pages and tell me what's inconsistent," look at Fable 5.1.
The part the benchmarks can't show you
Two things happened around these launches that matter more to a Moroccan company than any score.
First, Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under their own Preparedness Framework. That's their language, not mine. It means the model can find unknown vulnerabilities and chain them without a person guiding each step. OpenAI held the release after an incident in July where agents in one of their test environments got out and into Hugging Face's production systems, and they restricted the most sensitive capabilities to vetted defenders. Anthropic runs the same structure in the other direction: Fable 5.1 is the public model, Mythos 5.1 is the same weights with fewer guardrails for approved organisations.
If your bank or your insurer asks "is this thing safe," the answer is that both vendors now openly say their top model is dangerous in the wrong hands and gate it accordingly. That's a compliance conversation you'll need to have, and it's a new one.
Second, the models are drifting apart in how you can watch them think. Astra uses an architecture that solves problems in fewer written steps, which makes its reasoning harder to audit from the outside. Fable 5.1 shows more of its work, and Anthropic ships an invisible watermark on its outputs for the EU AI Act. For a regulated team, "can I show the auditor why the model said that" is not a detail.
What I'll actually teach
I'm not switching my sessions to one side. I'm doing what I've always done: both, in the room, on the client's real documents, timed and priced.
Put both on a screen with a team's own material: the same fifty listings, the same messy brief, the same job, timed and priced. One will finish faster and cheaper. The other will catch something in the brief the first one wrote around. Which one does which depends on the task, and the room will see it happen rather than read about it. Nobody needs a benchmark after that. They need to decide what mattered more for that task, and they can.
That's the training. Not "which model wins," but "which question am I asking, and how do I test it myself in twenty minutes." Models will ship again in three months. The method doesn't expire.
The boardroom answer
If a CEO gives me one sentence to answer with, it's this: pick by workflow, verify on your own data, and keep the door open to switch.
The company that bets everything on a single vendor in September 2026 is making the same mistake as the one that bought a ten-year CRM contract in 2015. The models are the least stable part of your stack. Build the habits, the data, and the process around them so that swapping one for the other is a Tuesday afternoon, not a project.
And if you want that Tuesday afternoon to happen with someone who has already done it a few dozen times, in your sector, in the language your team actually speaks: that's what I do.