Nate runs Anthropic's newest model (referred to throughout as "Fable 5.1") through real knowledge-work assignments instead of benchmarks: a GoPro/Starman acquisition DCF plus executive deck, a 100-word writing test, and a 37-second architectural walkthrough built from scratch in Blender using only a Seattle property address. The headline finding is that the model's low effort setting already produces genuinely usable work — a seven-sheet workbook and 13-slide deck with working formulas — while higher effort buys better reasoning hygiene rather than just more pages. He argues the real story is token efficiency and effort-tiering, not a model ranking, and recommends mixing models across passes.
- On the cheapest, fastest setting, the model researched the GoPro/Starman deal, built a seven-sheet Excel workbook with working formulas and scenarios, and produced a 13-slide deck landing on a $1.15/share base case — but omitted a sources sheet and a checks sheet, so the file was finished while the reasoning was hard to audit.
- At the highest effort setting it built a nine-sheet workbook and 15-slide deck, treated GoPro and Starman as separate businesses, added deal-close probability, a weighted average cost of capital, an exit-multiple check and 26 linked sources, arriving at roughly $1.30/share.
- The competing OpenAI model ("GPT 5.6 Soul") produced the most inspectable work — 10 sheets, a dedicated sources sheet and a checks sheet with an explicit pass result, base case $1.21 — but a less attractive deck.
- Nate's practical workflow: draft on low, use another model to check structure and add verification, then finish on high effort for polish — treating the first draft as a pass, not an endpoint.
- Knowledge work is hard for models because, unlike code, it never tells you when it's done; there's no compiler or test suite to signal correctness.
- In a 100-word Toyota history test, the new version dropped the older model's decorative metaphors ("thirsty Detroit models"), packed in more facts and a clearer causal chain; the OpenAI model traded dates for smoother narrative — better for general audiences versus executive ones.
- Given only an address, a loose brief and tool access, the model wrote Blender code to build the house, terrain, interiors, lighting, landscaping and camera path, rendered stills, inspected them, revised scenes, checked a motion preview and output a 37-second film — with no human touching Blender.
- Pricing is unchanged at $10/M input and $50/M output, but cache reads fell from $1 to $0.25 per million; Anthropic estimates ~25% lower cost on typical workloads and ~45% on highly agentic work, and Nate says it eats far less of his subscription limit than the prior version.
The sharpest framing in the video: a deliverable can look done while being impossible to verify, which is exactly the failure mode of AI knowledge work.
So, the file was finished, but maybe the thinking was hard to check... Now, the film might be the thing that you want to share virally. The spreadsheet is what helped me understand what Fable 5.1 is trying to do.
Concrete, itemized evidence of what a higher effort setting changes in the analysis rather than the output length.
It added a probability that the deal would close, made the funding need very explicit. It also used a weighted average cost of capital... and added an exit multiple check and linked 26 different sources. Extra didn't simply make a longer deck here. Extra actually found interesting questions that could change the investment decision.
You just don't need to take the Ferrari to the grocery store. Sometimes you're fine taking the Honda, and in this case it's a really nice Honda.
Code will tell a model when it is wrong... Knowledge work does not give you the courtesy of saying I am done.
With Fable 5.1, there's a difference between a complete file and supercomplete thinking.