AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Does Reducing The Astra Vs Fable Benchmark To Two Points Mean For AI Accuracy? on ThorstenMeyerAI.com

TL;DR

Recent revisions to the Astra vs Fable benchmark have significantly narrowed the score gap from five points to just two, raising questions about the validity of the metric. This development affects perceptions of AI intelligence and cost-efficiency, with ongoing uncertainty about the true performance of Astra.

Recent adjustments to the Artificial Analysis Intelligence Index have reduced the score gap between GPT-6 Astra and Fable 5.1 from five points to just two, according to sources familiar with the index’s latest revision. This change challenges earlier interpretations that Astra was significantly behind Fable in overall intelligence, and it underscores the importance of understanding the metrics behind these benchmarks. The revision impacts how AI performance and cost-efficiency are assessed, with implications for developers and users relying on these scores for decision-making.

The original comparison, widely circulated, claimed that Fable 5.1 scored 66 on the AI Index while Astra scored 61, suggesting a notable performance gap. However, the scores were based on an earlier version of the index, which was later revised to reflect updated evaluation methods and new metrics. In the current version, Fable 5.1 now scores approximately 57, and Astra scores around 55, effectively reducing the difference from five points to two. This change is not due to a decline in Astra’s capabilities but results from the index’s methodological updates, including the removal of certain metrics like GPQA Diamond and the addition of new ones such as AA-Briefcase and GDP.pdf.

Experts note that these score shifts highlight the volatility of benchmarking systems that are still evolving, especially as models and evaluation techniques become more complex. The revised scores suggest that Astra’s performance, while still slightly behind Fable, is much closer than initially reported. Importantly, the index’s own analysis indicates that Astra remains more cost-effective for coding tasks, but it does not outperform Fable in general intelligence efficiency, contradicting simplified narratives based solely on the raw scores.

At a glance
updateWhen: announced April 2024
The developmentBenchmark scores for Astra and Fable have been revised, reducing the previously reported five-point gap to just two points, prompting reevaluation of AI performance claims.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, „not a rounding error“ — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 „Five points is not a rounding error.“ Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. „140M vs 42M tokens“ compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and „largely sits behind its predecessor on the Intelligence Index vs cost frontier.“ Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. „Astra attacks the economics“ stretched a true coding result over an intelligence index where AA says the reverse.
04
„Max effort“ isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a „non-reasoning“ score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points „Same headline price“ hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and „Benchmarking GPT-6 Astra“ (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Revisions Reshape AI Performance Perception

The reduction in the Astra versus Fable score gap significantly alters the perception of Astra’s relative intelligence. While earlier figures suggested a clear lead for Fable, the updated scores imply that Astra is closer in performance than previously thought. This matters because many industry assessments, investor decisions, and competitive strategies depend on benchmark results. Furthermore, the revision exposes the fragility of current benchmarking methods, which can be affected by index updates and metric changes, potentially leading to misinterpretation of a model’s true capabilities.

For AI developers and users, understanding that scores are subject to change reinforces the importance of contextualizing benchmark results within the methodology and version history. It also raises questions about how much weight should be given to these scores when making strategic decisions or evaluating AI progress. Ultimately, the story shifts from a narrative of clear dominance to one of closer competition and the need for more robust evaluation frameworks.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evolving Benchmark Methodologies and AI Performance Metrics

The AI Index has undergone multiple revisions since its inception, reflecting ongoing efforts to improve the accuracy and relevance of its evaluations. The recent update, which coincided with Astra’s launch, involved replacing some metrics and recalibrating scoring baskets. Previously, the index appeared to favor models that externalized reasoning in tokens, which benefited architectures with latent or looped reasoning capabilities like Astra. As the index shifted to measure different aspects, scores for Astra and Fable changed accordingly.

This development underscores a broader challenge in AI benchmarking: models are rapidly evolving, and evaluation metrics often lag behind architectural innovations. Astra’s architecture, which reasons in latent space without explicitly verbalizing reasoning tokens, is not fully captured by token-based measures. The discrepancy between architecture and index measurement methods complicates direct comparisons and can lead to misleading conclusions about performance and efficiency.

Prior to these revisions, the narrative emphasized Astra’s cost-efficiency and its competitive edge in coding tasks. Now, with the scores closer, the focus shifts to understanding how architectural differences and evaluation methods influence the perceived performance gap, emphasizing the need for more comprehensive and architecture-aware benchmarks.

„The benchmark scores are a moving target, and relying on a single number without considering the index version and methodology is misleading.“

— Thorsten Meyer, AI researcher

Amazon

AI performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Impact of Architectural Changes on Scores

It remains uncertain how much the score revisions truly reflect changes in Astra’s capabilities versus methodological adjustments. Experts agree that Astra’s architecture, which reasons in latent space without explicit tokenized reasoning, is not fully captured by token-based benchmarks. Consequently, the true performance gap may be smaller or larger than the revised scores suggest. Additionally, the long-term stability of the benchmark and whether future revisions will further alter the scores remains unknown. OpenAI has not publicly detailed how these architectural innovations will be incorporated into standardized evaluation metrics, leaving some ambiguity about the future comparability of scores across models.

Amazon

AI model comparison charts

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Monitoring Benchmark Revisions and Architectural Developments

Going forward, industry analysts and researchers will closely monitor further updates to the AI Index and other benchmarking systems. There is a growing call for developing architecture-aware evaluation methods that accurately reflect models like Astra. OpenAI and other organizations may publish more detailed technical assessments of Astra’s architecture, clarifying how its reasoning process impacts performance metrics.

In addition, competition among AI developers is likely to shift focus from raw benchmark scores to more nuanced performance and efficiency measures. Expect further discussions on how to standardize evaluations that fairly compare models with different architectures and reasoning mechanisms. For Astra, the immediate next step involves transparency about how its architecture influences benchmark results and whether new metrics will better capture its strengths.

Amazon

AI accuracy testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the reduction in Astra’s score mean it is less capable?

No, the score change reflects revisions in the benchmarking index, not necessarily a decline in Astra’s capabilities. Its architecture may still outperform others in certain tasks, especially coding, but the scores are now more comparable and context-aware.

Why did the benchmark scores change after Astra’s launch?

The AI Index was updated to improve its evaluation methods, including replacing metrics and recalibrating scoring baskets. These revisions affected all models‘ scores, making previous comparisons outdated.

What does Astra’s architecture mean for its performance measurement?

Astra’s architecture reasons in latent space without explicitly verbalizing reasoning tokens, which makes token-based benchmarks less accurate in measuring its true performance. New evaluation methods may be needed to reflect its capabilities accurately.

Will future benchmarks continue to change?

It is likely, as benchmarking systems evolve to better capture architectural innovations. Stakeholders should interpret scores with awareness of the index version and methodology used.

How should I interpret Astra’s performance now?

With the revised scores, Astra appears closer to Fable than initially thought, especially in general intelligence metrics. However, its cost-efficiency in coding remains a notable strength, and architectural differences mean scores should be viewed as part of a broader performance context.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

14 AI Automation Tools To Power Smarter Workflows In 2026

Discover 14 top AI automation tools shaping smarter workflows in 2026, from agent builders to coding assistants, with insights on their applications and significance.

The 90-Day Window Closed. Nobody Sent a Notice.

The 90-day window for responsible disclosure has effectively ended, with no notices sent by vendors or researchers, raising security concerns.

One Founder, One Night, AI, 21 Packages: The Gewerkton Breakthrough

A solo founder built 21 verified software packages in one night using AI agents, creating Gewerkton, a construction documentation platform now in beta.

The 27% Problem: Why Google Wrote a $750M Check to Catch Anthropic

Google commits $750 million to boost enterprise AI with new platform, aiming to reclaim market share from Anthropic, which currently leads in enterprise AI adoption.