Claude Fable 5 for Finance: I Tested Anthropic’s New Model on Launch Day
There is a brand new Claude model that sits above Opus, and it is free on every paid plan until June 22nd. Anthropic dropped Claude Fable 5 the morning I wrote this, so I cleared my afternoon and put it through the same three tests I run on every model release. The point was simple: figure out exactly where it helps finance work before trusting it with anything real.
Model releases come fast, the announcements all sound the same, and you do not have a week to decide whether this one matters. Pick wrong in one direction and you are paying for capability you never use. Pick wrong in the other direction and you are handing month-end to a model that falls apart three steps into a complex task.
So here is what the new model is, how the access window works, what happened across my three tests on real finance files, and my honest take on who actually needs it.
The clock matters here: Fable 5 is included free on Claude Pro, Max, Team, and Enterprise plans through June 22. After that it drops off those plans and runs on usage credits until Anthropic adds capacity. You have a two-week window to test it without spending extra, which is exactly what this post helps you do.
What Is Claude Fable 5?
The naming is doing a lot of work, so a quick orientation. Fable 5 is the first public model from Anthropic’s Mythos class, a new tier that sits above Opus. Until launch day, Mythos-class capability was restricted to a small group of cybersecurity partners working with the US government. Fable 5 is that capability with guardrails added and made available to everyone.
Three things actually worth caring about as a finance person:
- It posted the highest score of any model on Hebbia’s Finance Benchmark, which tests senior-level reasoning over documents, charts, and tables. That is our world.
- Anthropic’s claim is that the longer and more complex the task, the bigger its lead over previous models.
- The access window. It is free on paid plans through June 22, then usage credits after that. Two weeks to test it for free.
You do not need to touch the API to try it. It is in the model picker at the bottom of the chat box right now, sitting next to Opus. Fable 5 is the most token-hungry model in the lineup, and the in-app notice tells you it draws down usage twice as fast as Opus. More on whether that is true after the tests.
Where Fable sits next to Opus, Sonnet, and Haiku
| Model | Tier | Best for | Cost reality |
|---|---|---|---|
| Fable 5 | Mythos class (new top tier) | Long, multi-step, multi-file work; senior-level reasoning | Free on paid plans through June 22, then usage credits; draws down usage faster per run than Opus |
| Opus 4.8 | Opus class | Heavy analysis and complex work, the previous ceiling | Included in paid plans |
| Sonnet 4.6 | Sonnet class | Everyday analysis, drafting, most day-to-day finance asks | Included in paid plans, cheap on API |
| Haiku 4.5 | Haiku class | Fast, high-volume, simple tasks | Cheapest |
The Opus 4.8 fallback, in one paragraph
Fable ships with safety classifiers. If a request trips one, Opus 4.8 answers instead and tells you it happened. According to Claude that lands in under 5% of sessions, and my finance work never came close. Across all three tests below, I was never once downgraded. If you see the notice mid-session, nothing is broken. You are getting the previous top model for that response.
My test setup: three finance tests, real files
I run every new model through the same dataset so results compare cleanly across generations. It is a three-location coffee shop business with the kind of files every finance team recognizes: a profit and loss statement, a point-of-sale file with roughly 150,000 transaction rows, a couple of mapping tables including a price list, prior variance commentary from location managers, and two days of bank data across multiple accounts per location.
The framework is always the same. Test one is data modeling: can it clean up and join messy files that do not share a clean key. Test two is reasoning: can it do the judgment work and come out with something defensible. Test three is autonomy: can it run a multi-step process from one prompt without me babysitting it.
Test 1: The data model job Power Query usually owns
This is the test I care most about, because it is the gap between a chatbot that summarizes a file and a tool that does analyst work.
Three files, no clean key
The setup is deliberately ugly. The point-of-sale file has 150,000 transactions with quantities but no dollar amounts. The P&L has monthly revenue in dollars but no transaction counts. The mapping tables sit off to the side with a price list and other reference data. The location and month formats do not match across files, the grains are different, and the metric I actually want, revenue per transaction, does not exist in any of them.
In Power Query this is the familiar grind: merge queries, fix the formats, group to monthly grain, build the relationship, write the measure. An hour if nothing fights you, and something always fights you.
The exact prompt
I went deliberately vague here, because Anthropic’s claim is that this model figures it out from a basic prompt.
I need to do some financial analysis. I have given you three data sets for the
F9 Finance coffee shop. I need to be able to do an in-depth analysis
understanding things like my cost per transaction and my average check, and none
of the data is living in the same files. Please give me the best analysis
possible and don't come back and ask me any questions. I want you to figure it
all out for yourself.
What came back
Fable inspected the files, worked out that the point-of-sale revenue tied to the GL to the penny, and identified the price list as the bridge between the two. Then it ran the full thing on its own: unit economics, margins by location, category and daypart mix, and a validation pass to make sure the blended data was trustworthy before it built anything on top.
The output came back as an Excel workbook with nine tabs. An executive summary, a unit economics tab with every per-transaction metric, category and daypart breakdowns, validation, and the underlying data tabs beneath all of it. It did the data blending, the cleaning, the validation, and the analysis, and it delivered more than I asked for. The formatting inside the artifact view was not something I would send as-is, but the substance was there. This is the work that normally takes an hour in Power Query, done from a vague prompt.
Test 2: A controller-grade review pass
Commentary vs. the numbers, where narratives fall apart
Every close has a review pass nobody automates. A location manager writes “labor was over because we added coverage,” and someone senior has to ask whether the traffic actually justified that coverage. Checking narrative against numbers is judgment work, and it usually lands on the most expensive person in the room at the most tired point of the month.
It is also where AI models have historically struggled. Give a model context that contradicts the data and it often takes the words as fact, or stalls out. So I fed Fable a CFO deck from a couple of months back and a commentary file from the location managers, some of which matches the financials and some of which directly contradicts them, and I planted five contradictions on purpose.
The review prompt
I have given you some prior commentary from our location managers as well as a
CFO deck we previously prepared. I need you to use this information to better
inform the financials. Note that I have not reviewed or validated this
information and I'm not sure what the quality level is. I need you to do a
complete analysis of the performance for our CFO including recommendations. And
I need you to do this without asking me any questions.
What the model found
Fable did not take the commentary at face value. Its own line was that the manager commentary was operationally useful but directionally untrustworthy. Of 18 claims it could test against the data, it flagged six as flat-out contradicted. I had only planted five, so it either found one I missed or counted slightly differently, but either way it did not miss the ones that mattered.
The report came back as a Word document, about four and a half pages, in a color scheme the tool picked on its own. It ran six sections: a reliability assessment of the inputs I gave it, a claim-by-claim commentary scorecard, the prior deck’s accuracy issues, first-half performance, a second-half forecast assessment, and recommendations. The conversion stumbled once inside the artifact view, but it opened cleanly in Word. This is the contradiction-handling that trips up most models, and Fable not only caught the gaps, it knew what to do about them.
Test 3: One prompt, an entire day of work
The task I refused to break up
This is the one Anthropic actually built the model for, and it is the one that changes how you use AI day to day. With every previous model, complex work meant chopping the task into pieces. Prompt, correct, prompt again. Five steps meant five prompts and three corrections, and you became the project manager of your own request.
So the third test was a full day of my job in a single prompt, with two days of bank data added on top of everything the model already had. Three real deliverables in three different formats, no clarifying questions allowed. And buried in the bank data is a problem I never mentioned: one account swings from a healthy balance into an overdraft. If it is really doing senior-level work, it finds that on its own.
I need you to replicate my entire job for the day. Start by getting a cash
position report to the treasurer and CFO, with a clear list of any changes that
need to happen. Then get out the entire month-end package as a PowerPoint for
the CFO and COO to understand what's happening in the business. Once that's
done, review the forecast from the teams and give me a challenger forecast for
the back half of the year, preferably inside a dynamic application so I can see
how my high-level assumptions compare to what the business is saying. Don't ask
me any clarifying questions.
What it shipped
Eight minutes. One prompt, no questions, never downgraded to Opus, and it came back with three finished deliverables in three formats.
The cash position report came back as a spreadsheet, and it caught the overdraft at the Astoria location on its own. It flagged it as a critical item and told me exactly what to do: call the bank immediately to attempt a wire recall, and run a same-day internal transfer from the Astoria reserve. That is the whole test, and it surfaced it unprompted.
The month-end package came back as a 13-slide PowerPoint, fully branded in my F9 template without me asking. Executive summary, cash position, cash and controls, location performance, and notes comparing the team forecast against the challenger view. Every slide built on the analysis from the earlier tests.
The challenger forecast came back as a working JavaScript application. Sliders for labor and cost of sales that move revenue up and down, a view of revenue against the team’s plan, and an explanation of why the two differ. All of it branded. I have had Claude build apps before, and I have never seen one come together like this from a single prompt as the last step in a chain.
The scorecard
| Test | What I asked | Result | Verdict vs. Opus |
|---|---|---|---|
| Data model | Cross-file join, no clean key, full analysis from a vague prompt | 9-tab Excel workbook with validation, tied to GL to the penny | ~20-30% better. Opus could integrate the files, but likely not all the way to a Power Query replacement |
| Review pass | Commentary vs. financials with contradictions planted | 4.5-page Word report, 6 of 18 claims flagged as contradicted | ~20-30% better. Opus catches contradictions but often does not know how to resolve them |
| Autonomy | A full day of work in one prompt, no questions | Cash report, 13-slide deck, and a working forecast app in 8 minutes; caught the overdraft unprompted | ~50-60% better. Each layer built on the last and it behaved like an analyst |
Does Fable live up to the hype, and is it substantially better than Opus? Yes, and I agree with Anthropic that the gap grows the bigger the task gets. The third test is the one that sold me. It did an entire day of work, every layer building on the one before it, in roughly the time a single one of those deliverables would have taken Opus.
Where it falls short, and the honest verdict
A few real knocks. The way Fable describes its own work is rough. It reaches for the biggest words it can find, and where Sonnet and Opus narrate what they are doing in plain language, Fable sounds like it is studying for the SAT. It does not affect the output, but it is a step back in polish. The artifact formatting was not something I would forward as-is, and one document conversion failed and had to be opened in Word.
On tokens, the in-app notice says Fable draws down usage twice as fast as Opus. In my testing it was not close to twice. It is noticeably bigger than Opus per run and maybe around twice Sonnet, but it also gets to the answer faster and skips the prompt-and-reprompt cycles entirely. By the time the whole job is done, you are probably spending only 10 to 20% more tokens than you would have, not double.
The June 22 cliff is the thing to actually plan around. If Fable becomes part of your workflow over the next two weeks, expect it to cost usage credits after that date until Anthropic restores it to subscription plans. Do not build a month-end process around a model you might lose access to mid-close.
And not everyone needs to turn this on. If your AI use is drafting emails to an auditor and explaining a DAX measure, Sonnet is still the right answer and the cheap one. The new tier earns its keep on long, messy, multi-file work: the data models, the review passes, the multi-step processes. That is also exactly the work that eats your evenings, which is why this release matters more for finance than the last several. Matching the model to the task is the skill.
How to run this test on your own files this week
You do not need my dataset. You need 30 minutes and files you already have.
- Open Claude on any paid plan and select Fable 5 in the model picker at the bottom of the chat box. The free window runs through June 22.
- Pick two files from your own world that do not share a clean key. A system extract and a manual tracker. A trial balance and a headcount file. Every team has a pair like this.
- Run the join test. Ask for a metric that needs both files and a vague-on-purpose prompt, then tell it to figure it out without asking you questions. See whether it finds the bridge on its own.
- Run the review test. Give it your last close’s commentary and the matching actuals, tell it you have not validated the inputs, and ask it to analyze performance for your CFO. Compare its pushback to what your actual review caught.
- Run the autonomy test. Write one prompt with every step of a small recurring process, ban clarifying questions, and see how far it gets. Count how many corrections you would have needed with your current model.
- Judge it on your files, not benchmarks. If the new tier does not beat your current setup on your real work, the upgrade can wait. If it does, you will know exactly where it pays for itself.
The model picker is the cheapest experiment in finance right now. For the next two weeks it is free.

thanks for info.