top of page
Search

Claude Fable 5.1: Features, Benchmarks, Pricing, and What It Means for AI Agents

Philip Moses
Sep 9
5 min read

AI models are increasingly moving beyond simple conversations and into long-running, autonomous workflows. Anthropic’s Claude Fable 5.1 is built around this shift, focusing on coding, research, automation, and tasks that may require an AI agent to work for hours.


In this blog, we’ll look at what makes Fable 5.1 different, its key features, benchmark performance, pricing, how it compares with other Claude models and GPT-5.6 Sol, and where it makes the most sense to use it.

What Is Claude Fable 5.1?

Claude Fable 5.1 is Anthropic’s latest frontier model designed for long-horizon coding, research, and agentic work.

Unlike models primarily optimized for everyday conversations, Fable 5.1 is designed to take on larger objectives, plan multiple steps, use tools, verify its work, and continue working with less human intervention.


Key specifications include:

  • 1 million token context window

  • 128K maximum output

  • $10 per million input tokens

  • $50 per million output tokens

  • $0.25 per million cached tokens read

  • Adaptive thinking with Low, Medium, and High effort levels

  • Knowledge cutoff of June 2026

One of the most important changes is the cost of cached context. Fable 5.1 reduces cache-read pricing by 75% compared with Fable 5, making repeated context-heavy workloads much more affordable.


Anthropic estimates that typical workloads can be around 25% cheaper, while highly agentic workloads can see savings of up to approximately 45%.

What Makes Fable 5.1 Different?

The biggest difference is its focus on long-running AI agents.


Instead of asking an AI to complete one small task at a time, developers can give Fable 5.1 a broader objective and allow it to work through multiple stages.


For example, an agent could:

  • Understand a project requirement

  • Research the existing codebase

  • Create an implementation plan

  • Write and test code

  • Identify problems

  • Correct its approach

  • Continue working on the remaining tasks

This makes Fable 5.1 particularly interesting for software engineering, scientific research, debugging, and autonomous development workflows.


Anthropic has highlighted examples of Fable 5.1 being used for multi-hour workloads, including an reported 38-hour unattended machine-learning run involving experiments, analysis, and corrections.


This points toward an important change in how AI assistants may be used: from answering questions to completing projects.

Fable 5.1 Performance on Agentic Tasks

One of the strongest results for Fable 5.1 comes from Terminal-Bench-Science 0.1, which evaluates agentic scientific research performed through a terminal.


Fable 5.1 achieved a score of 52.6%, compared with:

  • Claude Opus 5 — 29.0%

  • Claude Fable 5 — 24.7%

  • GPT-5.6 Sol — 22.4%


This means Fable 5.1 achieved more than twice the score of Fable 5 on this particular benchmark.


The result is particularly significant because these tasks require more than generating a correct response. The model needs to reason through a problem, interact with tools, execute steps, investigate results, and adapt its approach.

Claude Fable 5.1 Benchmark Results

Anthropic’s published comparison shows Fable 5.1 performing strongly across coding, automation, knowledge work, reasoning, and computer-use benchmarks.


Benchmark

Fable 5.1

Fable 5

Opus 5

GPT-5.6 Sol

Terminal-Bench-Science 0.1

52.6%

24.7%

29.0%

22.4%

Terminal-Bench 4.0

55.8%

42.0%

52.3%

37.3%

CursorBench 3.2.0

73.4%

70.5%

70.0%

67.2%

GDPval-AA v2 (Elo)

1853

1723

1824

1711

AutomationBench

31.4%

17.1%

26.9%

19.6%

Humanity's Last Exam

65.0%

63.8%

63.6%

Not published

OSWorld 2.0 (strict)

41.7%

36.1%

39.6%

Not published


According to Anthropic, these figures were measured with Fable 5.1’s production safeguards enabled.

Stronger Coding Performance

Fable 5.1 shows particularly strong results on coding and terminal-based tasks.

On Terminal-Bench 4.0, Fable 5.1 achieved 55.8%, ahead of Opus 5 at 52.3% and Fable 5 at 42.0%.


It also achieved 73.4% on CursorBench 3.2.0, compared with 70.5% for Fable 5 and 70.0% for Opus 5.


These results suggest that Fable 5.1 is particularly suited to software-development workflows where the model needs to work through multiple steps rather than simply generate code.

Knowledge Work and Automation

Fable 5.1 also performs well on business-oriented tasks.

On GDPval-AA v2, it achieved an Elo score of 1853, compared with 1824 for Opus 5 and 1723 for Fable 5.


On AutomationBench, Fable 5.1 scored 31.4%, compared with 26.9% for Opus 5 and 17.1% for Fable 5.


This matters because business AI is increasingly moving toward workflows where models need to interact with systems, follow processes, and complete several actions rather than simply produce text.

Scientific Research and Vulnerability Analysis

Another important area for Fable 5.1 is research.


Anthropic reports that the model was used to create a high-resolution elevation map covering a portion of Venus using radar imagery collected by NASA’s Magellan mission.


The work reportedly improved feature resolution from approximately 10–20 kilometers to 2–3 kilometers, while also improving height accuracy.


Fable 5.1 can also be used for software vulnerability research, although Anthropic maintains restrictions around exploit development and other high-risk cybersecurity activities.


The company says its newer safeguards reduce false positives, allowing legitimate security research to be carried out with fewer unnecessary interventions.

Claude Mythos 5.1: A More Restricted Model Variant

Alongside Fable 5.1, Anthropic has introduced Claude Mythos 5.1.

Mythos 5.1 uses the same underlying model but operates with different safeguards for approved cybersecurity and life-sciences research.


Access is restricted through Anthropic’s trusted programs.


The distinction is important because Mythos 5.1 is not simply a more powerful model. Instead, the primary difference is what the safeguards allow the model to do.


Anthropic reports Mythos 5.1 achieving 60.9% on Terminal-Bench 4.0, compared with 55.8% for Fable 5.1.

Claude Fable 5.1 Pricing

Fable 5.1 keeps the same list pricing as Fable 5.

Pricing

Cost per Million Tokens

Input

$10

Output

$50

Cache read

$0.25

Cache write — 5 minutes

$12.50

Cache write — 1 hour

$20

Batch API

50% discount on input/output

The 75% reduction in cache-read pricing is one of the most important changes.


For long-running agents that repeatedly access the same code, documentation, or instructions, cheaper cached context can significantly reduce the overall cost of an AI workflow.

Which Claude Model Should You Choose?

Fable 5.1 is not necessarily the best choice for every task.

Use Case

Recommended Model

Why

Multi-hour autonomous agents

Fable 5.1

Designed for long-running agentic work

Thorough repository planning

Opus 5

Strong reasoning at a lower price

Everyday coding and chat

Sonnet 5

Faster and more affordable

High-volume, latency-sensitive work

Haiku 4.5

Lower cost and faster responses

Approved advanced security/science work

Mythos 5.1

Designed for specialized workflows

Cost-sensitive Fable workloads

Fable 5.1 Low/Medium effort

Can reduce reasoning costs

For everyday conversations and quick coding tasks, Sonnet 5 may be the more practical option.


For complex reasoning and repository planning, Opus 5 remains attractive.

But for long-running autonomous work, Fable 5.1 is where Anthropic is positioning its strongest capabilities.

Why Fable 5.1 Matters for Businesses

The most important takeaway from Fable 5.1 may not be any individual benchmark.

It is the broader direction of AI development.


Businesses are increasingly looking for AI systems that can do more than generate content. They want systems that can:

  • Analyze large amounts of information

  • Work with existing software

  • Automate repetitive workflows

  • Conduct research

  • Write and test code

  • Monitor progress

  • Correct mistakes

  • Complete tasks with less supervision

Fable 5.1 is designed around this model of AI.

Its large context window, adaptive reasoning, lower cache costs, and long-running agent capabilities make it particularly relevant for organizations exploring autonomous AI systems.

Final Thoughts

Claude Fable 5.1 represents a shift from AI that answers questions to AI that can work on problems for extended periods.

Its strongest improvements are visible in agentic science, coding, automation, and long-running workflows. At the same time, the model keeps Fable 5’s list pricing while significantly reducing the cost of cached context.

The key takeaway is simple:

Fable 5.1 is not primarily designed to win every chat. It is designed to make long-running AI work more capable and more practical.

For everyday chat and coding, smaller Claude models can still offer a better balance of speed and cost. But when an AI agent needs to plan, execute, test, adapt, and continue working for hours, Fable 5.1 is built for that future.

 
 
 

Recent Posts

See All

Comments


Curious about AI Agent?
bottom of page