Saturday, 5 September 2026
Nate runs Anthropic's newest model (referred to throughout as "Fable 5.1") through real knowledge-work assignments instead of benchmarks: a GoPro/Starman acquisition DCF plus executive deck, a 100-word writing test, and a 37-second architectural walkthrough built from scratch in Blender using only a Seattle property address. The headline finding is that the model's low effort setting already produces genuinely usable work — a seven-sheet workbook and 13-slide deck with working formulas — while higher effort buys better reasoning hygiene rather than just more pages. He argues the real story is token efficiency and effort-tiering, not a model ranking, and recommends mixing models across passes.
- On the cheapest, fastest setting, the model researched the GoPro/Starman deal, built a seven-sheet Excel workbook with working formulas and scenarios, and produced a 13-slide deck landing on a $1.15/share base case — but omitted a sources sheet and a checks sheet, so the file was finished while the reasoning was hard to audit.
- At the highest effort setting it built a nine-sheet workbook and 15-slide deck, treated GoPro and Starman as separate businesses, added deal-close probability, a weighted average cost of capital, an exit-multiple check and 26 linked sources, arriving at roughly $1.30/share.
- The competing OpenAI model ("GPT 5.6 Soul") produced the most inspectable work — 10 sheets, a dedicated sources sheet and a checks sheet with an explicit pass result, base case $1.21 — but a less attractive deck.
- Nate's practical workflow: draft on low, use another model to check structure and add verification, then finish on high effort for polish — treating the first draft as a pass, not an endpoint.
- Knowledge work is hard for models because, unlike code, it never tells you when it's done; there's no compiler or test suite to signal correctness.
- In a 100-word Toyota history test, the new version dropped the older model's decorative metaphors ("thirsty Detroit models"), packed in more facts and a clearer causal chain; the OpenAI model traded dates for smoother narrative — better for general audiences versus executive ones.
- Given only an address, a loose brief and tool access, the model wrote Blender code to build the house, terrain, interiors, lighting, landscaping and camera path, rendered stills, inspected them, revised scenes, checked a motion preview and output a 37-second film — with no human touching Blender.
- Pricing is unchanged at $10/M input and $50/M output, but cache reads fell from $1 to $0.25 per million; Anthropic estimates ~25% lower cost on typical workloads and ~45% on highly agentic work, and Nate says it eats far less of his subscription limit than the prior version.
The sharpest framing in the video: a deliverable can look done while being impossible to verify, which is exactly the failure mode of AI knowledge work.
So, the file was finished, but maybe the thinking was hard to check... Now, the film might be the thing that you want to share virally. The spreadsheet is what helped me understand what Fable 5.1 is trying to do.
Concrete, itemized evidence of what a higher effort setting changes in the analysis rather than the output length.
It added a probability that the deal would close, made the funding need very explicit. It also used a weighted average cost of capital... and added an exit multiple check and linked 26 different sources. Extra didn't simply make a longer deck here. Extra actually found interesting questions that could change the investment decision.
You just don't need to take the Ferrari to the grocery store. Sometimes you're fine taking the Honda, and in this case it's a really nice Honda.
Code will tell a model when it is wrong... Knowledge work does not give you the courtesy of saying I am done.
With Fable 5.1, there's a difference between a complete file and supercomplete thinking.
A head-to-head of OpenAI's "GPT-6 Astra" (run through Codex) against Anthropic's "Claude Fable 5.1" across benchmarks and four one-shot practical tests: a browser Fortnite clone, an AI travel landing page, a 15-second motion-graphics explainer, and a 3D globe flight dashboard. Astra won three of the four and tied on motion graphics, while also being roughly half the token cost for similar benchmark scores. The conclusion isn't that Fable 5.1 is bad, but that Anthropic's usage-limit practices plus Astra's edge make splitting subscriptions between both worth considering.
- On reported benchmarks GPT-6 Astra beats Fable 5.1 on essentially everything, but the more meaningful gap is cost: on Terminal Bench 4.0 both land around 56% accuracy while Astra costs $10.35 versus $19.50 for Fable 5.1.
- In the one-shot Fortnite clone test, Codex/Astra produced a playable battle-royale with glider, chests, inventory, building and a map in about 45 minutes; Fable took roughly 90 minutes and 750,000 tokens and came out janky with shaky camera, wonky animations and bad gunplay.
- For the AI travel landing page, Astra used its built-in image model to produce a clean site that didn't read as AI slop, while Fable 5.1 returned a generic light-blue-to-dark-blue layout with self-generated graphics and no motion — described as "night and day."
- The presenter argues front-end results depend heavily on the operator: skilled users supplying reference images and component libraries make the base model matter less, but for median "just build me this" prompting Astra clearly wins.
- Motion graphics was a genuine tie — both models called the Higgsfield MCP and the same skill, routed to Seedance 2.5, and returned solid 15-second explainers.
- On the 3D globe dashboard, Astra's "Orbit" was clean and near-usable with a golden-hour sun-chasing feature, while Fable's "Arclight" prioritized visual spectacle to the point of being hard to read and "many prompts away" from usable.
- Anthropic reported far fewer benchmarks than OpenAI, which the presenter flags as a bummer for comparison.
- Practical recommendation: instead of a $200 20x Anthropic plan (where 20x isn't really 20x and weekly usage is halved), consider two $100 5x plans, one with each provider, to test both on real work.
The single most concrete number in the video — near-identical performance at roughly half the token cost.
At max, we're getting 55.8% accuracy with Fable 5.1, and with Astra, it's 56.7... But where they aren't the same is the cost. At max, I'm at $10.35 for Astra, and I'm at $19.50 with Fable 5.1.
Direct side-by-side of time, token cost and output quality on the hardest task in the test set.
For reference, this took Claude about like an hour and a half to create this, and about 750,000 tokens... overall, I would say just feels a little less polished.
We can say GPT-6 is a big leap forward on the OpenAI side, and Fable 5.1 is just giving us more of what we already like.
It almost feels like visual spectacle became the number one priority with Fable 5.1 versus functionality.
No video, no test, no benchmark is really going to be the same as when you get in there and use it yourself for your projects.