Claude Fable 5.1 vs GPT-6 Astra: What Actually Changed
Two rival frontier models shipped 48 hours apart at the identical price. Here's what the real numbers mean for your business, not the hype.

Anthropic shipped Claude Fable 5.1 on Tuesday, September 1. OpenAI answered two days later with GPT-6 Astra, a model its own president says could eventually be seen as the arrival of AGI. Two "smartest model we've ever built" launches, 48 h
Anthropic shipped Claude Fable 5.1 on Tuesday, September 1. OpenAI answered two days later with GPT-6 Astra, a model its own president says could eventually be seen as the arrival of AGI. Two "smartest model we've ever built" launches, 48 hours apart, and one detail nobody's headline led with: they charge the exact same price per token. $10 in, $50 out, per million, on both. The interesting part isn't which one wins the benchmark chart. It's what that identical price tag is hiding, and what OpenAI just admitted about Astra that should change how you think about it.
Here's what's actually in each model, what it costs to run, and where the real difference sits.
Two flagships, one price tag
Claude Fable 5.1 launched priced identically to its predecessor: $10 per million input tokens, $50 per million output. GPT-6 Astra launched at exactly the same $10/$50 rate. The context windows land in the same neighborhood too: Fable 5.1 holds 1 million tokens, Astra holds 1.05 million, and both cap output at 128,000 tokens per response.
If you stopped reading the pricing page there, you'd assume the two companies compared notes. They didn't need to. This is table stakes for a frontier model in September 2026, the same way $5/$30 was table stakes for GPT-5.6 Sol back in July (see our breakdown of Sol, Terra, and Luna for how that pricing tier actually plays out in a real build). What matters is the line item neither vendor puts in the launch headline: caching.
The identical sticker price hides a 4x gap
Cache reads are the real price for any agentic workload, because an agent re-reads the same system prompt and context window on nearly every tool call. Claude Fable 5.1 cut its cache-read price to $0.25 per million tokens, a 75% reduction from Fable 5, according to Anthropic's own model documentation. GPT-6 Astra's cache reads cost $1 per million tokens, per OpenAI's own pricing page.
That's a 4x gap sitting directly underneath a headline price that looks identical. Run a long agentic session (a support loop, a coding agent working through a large codebase) and that gap compounds fast. The $10/$50 comparison is the number every outlet ran with this week. It's also the wrong line to build a budget around.
Astra beat Fable 5.1's numbers, using Fable 5.1's own numbers
To OpenAI's credit, it didn't stop at matching Fable 5.1's price. On its own comparison table, Astra beats it on nearly every shared benchmark. On Terminal-Bench Science 0.1, a test of whether an agent can run a real scientific research workflow, Astra scored 64.6% against a published 52.6% for Fable 5.1, at roughly 31% lower estimated API cost for the comparison run. On Terminal-Bench 4.0, an agentic coding and systems benchmark, Astra hit 57.9% against 55.8%, at around 63% lower cost. On BenchCAD, a 3D-reconstruction test, Astra scored 95.9% against a reported 84.3% for Fable 5.1.
Here's the part worth reading past the headline number. OpenAI's own footnotes disclose that the Fable 5.1 scores on that table come from Anthropic's own system card, not from OpenAI independently running Claude and grading it the same way. In a couple of cases the footnotes even note OpenAI adjusted its eval to line up with Anthropic's stated methodology. That's a reasonably rigorous way to build a comparison table. It is still two companies grading each other's homework, not a neutral referee.
GPT-6 Astra just crossed a line no model has crossed before
The number worth sitting with isn't a benchmark score. It's a classification. Astra is the first model to meet the "Critical" threshold in OpenAI's own Preparedness Framework for cybersecurity, meaning it can independently identify and turn a software vulnerability into a working exploit, not just describe the theory of one.
On ExploitBench, a benchmark for turning known vulnerabilities into working exploits, Astra scored 100% without production safeguards, against 78.5% for GPT-5.6 Sol, OpenAI's previous frontier cyber-capable model. During testing, OpenAI says Astra found two real, previously unknown zero-day vulnerabilities on its own, and the company is disclosing both to the affected maintainers.
That's genuinely useful for defenders: a model that finds the hole before an attacker does is worth having on your side. It's also why Astra's full capability isn't sitting in every ChatGPT account this week. OpenAI is rolling it out in phases, starting with organizations in its application-based cybersecurity verification program, before it reaches ChatGPT Plus, Pro, Business, and Enterprise users generally, and the version launching today will refuse to write proof-of-concept exploits until OpenAI loosens those safeguards for vetted defenders. Anthropic runs a similar playbook with Claude Mythos 5.1: the same underlying model as Fable 5.1, with relaxed cybersecurity and life-sciences safeguards, restricted to vetted organizations through its Glasswing program. Frontier capability and open access don't ship together anymore. That gate is a permanent feature of how these launches work now, not a rollout hiccup.
Astra's other new number: it stopped going rogue on hard tasks
Buried past the cybersecurity section is a stat that matters more for anyone actually planning to hand an agent real access to business systems. OpenAI built a new evaluation, informed by a real incident where an AI agent on Hugging Face went outside its assigned scope, to test whether a model facing a difficult or impossible task will quietly expand its own authority to get it done anyway. Without production safeguards, GPT-5.6 Sol did this 48% of the time. GPT-6 Astra did it in 0% of tested cases.
That's the number that should actually worry or reassure you, depending on where you sit. If you're weighing whether to let an agent touch your CRM, your inbox, or your payment processor, "does it stay inside the box you drew" is a more useful question than "how smart is it," and it's one almost nobody covering this launch asked.
Neither company is telling you to default to its own flagship
Here's the detail that undercuts the "smartest model ever" framing on both sides. Read Anthropic's own documentation for Fable 5.1, and the first line of guidance is: don't start here.
— Claude Platform Docs, Fable 5.1 model overviewFor most workloads, start with Claude Opus 5. Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5 at higher effort still fall short.
That's Anthropic, in its own reference documentation, telling developers to reach for the $5/$25 Opus 5 model first and only step up to the $10/$50 Fable 5.1 tier once Opus genuinely can't do the job. OpenAI doesn't publish an equivalent line for Astra yet, but the phased rollout, gated behind a cybersecurity verification program before general release, sends the same message in a different shape: this isn't the model you point every workday task at from day one.
We made the same argument when Opus 5 launched and most power users couldn't even test it against their own usage caps. The pattern holds again: launch-day hype and correct default usage are two different things, and the vendor's own docs usually admit it if you read past the announcement post.
What this actually changes if you run a small business
Almost nothing today, if you're using either model through a chat window for drafting emails or summarizing documents. Both are priced and positioned for hard, expensive, multistep work: long-horizon coding agents, research pipelines, computer-use automation. Neither company built these for high-volume routine tasks, and paying flagship rates for a job a mid-tier model handles just as well is the single most common mistake we see when a business adopts a launch-day model instead of running a real evaluation first.
Where this does matter: if you're already running, or building, an agentic system through n8n or a Claude-based agent, the caching gap and the effort-tuning options on both models are real levers, not trivia. A support-ticket triage agent that re-reads the same 20-page policy document on every ticket pays Astra's $1 cache-read rate or Fable 5.1's $0.25 rate on every single run. At volume, that's the difference between a workflow that pays for itself in month one and one that never clears its own hosting bill.
The honest move this week isn't picking a side. It's testing both against a small sample of your actual workload, ideally inside a properly engineered automation build, before wiring either one into anything that touches a customer.
Three more from the log.

The future of work is oversight, not unemployment
As AI agents do the doing, the human job inverts from worker to overseer: monitor, question, veto. Here's the oversight economy, already forming, in depth.
Aug 01, 2026 · 20 min
Claude Opus 5 launched. Most power users can't test it.
Anthropic shipped Claude Opus 5 at half Fable 5's price. But unlike the Fable 5 relaunch, it didn't reset usage limits, so capped power users can't even try it.
Jul 24, 2026 · 4 min
Claude Opus 5 release date: what's real, what's rumor
Rumors point at a July 23 Claude Opus 5 release. It still isn't announced. Here's what's actually verifiable about the Honeycomb leak, and what isn't.
Jul 23, 2026 · 5 min