vibehacker
Discuss
Amara Nwosu
2 days ago

36 hours on Fable 5.1 in Claude Code: the cache-read price cut matters more than the benchmarks

Claude Code
An AI coding agent for terminal, IDE, web, and Slack

Anthropic put out Fable 5.1 and Mythos 5.1 on Monday: https://www.anthropic.com/claude-fable-and-mythos-5-1

Same model, different safeguards. Fable is the one you and I get; Mythos is behind a trusted-access programme for cyber and bio work.

I switched our two-person shop to Fable 5.1 in Claude Code on Monday night. Not a benchmark, just the KYC repo and our usual backlog. Notes:

  • The post claims "an estimated 25% less than Fable 5 for typical workloads", coming from cheaper cache reads, and "up to approximately 45%" for agentic work. Our sessions are nearly all agentic. The bill graph agrees so far. That is the headline for a small company, not the Terminal-Bench number.
  • It defaults to High effort in Claude Code. You can turn that down. I did for the boring tickets and it was fine.
  • Long sessions stay readable. One of their quotes says "prior models became hard to follow the longer they worked" and that matched my experience with Fable 5. Fewer "wait, why did it do that" moments in a two-hour run.
  • Zero data retention is available to "eligible customers" now, with the customer-controlled storage thing (EFS) coming "later this fall". I'll believe the fall part when I see it, but ZDR now is enough for our compliance guy.

Their own table has Terminal-Bench 4.0 at 55.8% vs 42.0% for Fable 5. I have no way to check that. I can check my bill.

3

3 comments

Join the discussion

Log in to comment.

  • "An estimated 25%" for "typical workloads" is doing a lot of work in that sentence. If your prompts are not cache-friendly you get nothing. Not saying it is wrong, saying read the small print before you tell your CFO.

    • Amara Nwosu

      Fair. Ours are cache-friendly by accident, because Claude Code reuses the same system prompt and the same files all session. If you are hitting the API raw with fresh context every call, you are the "typical workload" they are estimating for, and you should measure it yourself.

  • Jonas Kessler

    The safeguards paragraph deserves attention too: they state Fable 5.1 "can now be used to discover software vulnerabilities, though not to develop exploits for them", and that the cyber safeguards produce 60% fewer false positives. For those of us who were getting refused on ordinary security review tasks, that is a more practical change than any benchmark row. I have not verified the 60%; I have noticed fewer refusals this week.