The Rules of the LLM War Have Changed — How Should Ordinary People Choose?

The Rules of the LLM War Have Changed — How Should Ordinary People Choose?

On September 1st local time, Anthropic released its new flagship, Fable 5.1, and on the same day made a version with a different safety tier, Mythos 5.1, available to vetted institutions.

Convention would suggest that a launch at this scale should lead with benchmark scores. This time, it didn’t. Half the headlines across global tech media were about price cuts. Cache read costs dropped 75%. The expected cost of a typical task dropped 25%. Highly automated agent tasks saw cost drop by as much as 45%.

One of the most capable models in the world made price cuts its headline feature.

A top-tier model maker spending its flagship launch talking about price is a bit like Porsche holding a press event and skipping 0-to-60 numbers to talk about fuel economy instead.

This isn’t an ordinary version update. When a leading lab starts building its core narrative around price, it means the labs themselves know that benchmark scores alone can no longer create meaningful differentiation. The rules of the LLM war have changed.

Below, two things get unpacked in full: how this fight arrived at price as its central battleground, and how an ordinary person should actually choose among the flood of available models.

Every Leap in Capability Has Run on a Different Fuel

The evolution of large language models is, at bottom, a history of fuel changes. Every leap in capability has come from burning a different scarce resource.

It started with the pretraining era, burning compute and public text corpora. The leap from GPT-3 to GPT-4’s generation came from stacking parameters, data, and compute together — models chewed through nearly the entire public text internet. Competition in that phase was straightforward: whoever could afford more GPUs won.

Once stacking raw material stopped moving the needle, the industry entered the alignment era, and the fuel switched to human feedback data. Models learned to communicate well through techniques like RLHF, whose raw material is human annotation and preference ranking, one example at a time. The same base model, fed different feedback, produced wildly different usability.

Now the industry has entered a third phase. Models are no longer satisfied with answering questions — they’re starting to do the work themselves, and the fuel has switched again, this time to data from real-world tasks.

Fable 5.1, the model released this week, represents this phase. On Terminal-Bench-Science 0.1, an agentic research benchmark, Fable 5.1 scored 52.6%, up from 24.7% for the previous generation, Fable 5 — more than double in a single year.

Benchmark numbers feel abstract, so here’s a concrete example. Investment firm Millennium had a piece of internal code that crashed roughly once every million runs. Engineers had been chasing the cause for four or five years, and Fable 5 couldn’t crack it either. Fable 5.1 compared the disassembled output of an external library against core dump files layer by layer, and eventually traced the problem to a hidden defect buried inside that vendor’s library.

This kind of capability can’t be built from static text corpora. A model has to have seen a massive number of genuine execution traces, hit real dead ends, before it knows where to even start looking.

Line the three phases up together, and a pattern emerges. Compute can be bought — Anthropic signed a $35 billion compute deal with Lambda the same day, one of the largest cloud deals in AI history; money solves that problem. Algorithms spread rapidly through the open-source ecosystem, and the technical gap between leading labs keeps shrinking. Data is the one thing that’s getting harder and harder to buy.

Public internet text has been scrubbed and reused repeatedly, and the incremental supply has essentially peaked. What will actually separate the leaders in the next phase is specialized domain data, real interaction data, and proprietary enterprise data.

One detail worth sitting with: on the same day Fable 5.1 launched, Anthropic announced a new enterprise-grade data protection scheme, keeping logs and keys on the customer’s own cloud — a scheme designed in collaboration with more than a hundred enterprises. A leading lab voluntarily handing data sovereignty back to customers means everyone is implicitly agreeing on one thing: the ownership and circulation of data is being repriced.

Confirming ownership of data, pricing it, and trading it are becoming infrastructure-level problems. Some projects are already experimenting with on-chain identity to make individual data contributions provable and priceable — the kind of thing that sounds like an abstract concept in normal times, but starts to look like it was built for exactly this moment once you set it against a backdrop of data scarcity.

The second half of the model war isn’t about who has more compute anymore. It’s about who has more ammunition.

How Strong Are Today’s Models, Really?

The answer might be a little counterintuitive: the gap between leading models has already narrowed to the point where ordinary users can’t perceive it.

On GDPval-AA v2, a comprehensive knowledge-work benchmark, Fable 5.1 scored 1853, its same-generation sibling Opus 5 scored 1824, and the prior generation Fable 5 scored 1723. The gap between first and third place is under 2% — and that lead sits within the range of statistical noise. Two years ago, a version update meant a generational leap. Now what you get is a fight over decimal points.

And this upgrade isn’t uniformly better across the board, either.

Third-party testing found that Fable 5.1, running at its highest reasoning intensity, costs an average of $3.76 to complete a task — 20% more expensive than the previous generation, because output token volume reached 1.7x the prior generation’s, and the savings from caching didn’t fully offset the added output cost. Code review platform CodeRabbit ran its own tests: across 45 review tasks, Fable 5.1 found roughly the same number of issues as its predecessor, but the number of comments dropped from 253 to 166, with trivial nitpicks cut from 265 down to 79 — at the cost of average review time rising from 12.5 minutes to 18.5 minutes, nearly 50% slower.

Say less, think more — that’s what this upgrade actually looks like. The ceiling on capability has genuinely risen, but the trade-off between money, time, and quality hasn’t gone anywhere.

The labs clearly see this too, which is why, once benchmark competition stopped moving the needle, the battlefield shifted to two things. One is cost-efficiency — this round’s 75% cache price cut is aimed directly at the pain point of agents repeatedly re-reading context. The other is stability on long-running tasks — running for hours without errors or drift is the genuinely scarce quality in the automation era.

This is exactly why the strongest model on the market put price in its launch headline. The tail end of the benchmark era is the beginning of the application era.

How Should Ordinary People Actually Choose?

A basic starting point: for the vast majority of people, the bottleneck was never that the model wasn’t capable enough — it’s that they hadn’t thought clearly about what they actually need the model to do. Choosing a model doesn’t require chasing the strongest option. It just requires aligning three things: task type, cost sensitivity, and privacy requirements.

The Arena platform maintains an Agent task leaderboard, scoring models by their overall performance on real agentic tasks — a much closer proxy for actual work than a pure benchmark leaderboard. Fable 5.1 is too new to appear on it yet; the current leaderboard’s top spots go to Claude Opus 5 (High, 13.74%), Claude Opus 5 (Max, 11.69%), Claude Fable 5 (High, 10.61%), and GPT-5.6 Sol (xHigh, 9.49%). Sixth place belongs to Moonshot AI’s Kimi K3 (Max, 8.71%).

That sixth-place finish for Kimi K3 deserves a mention. A Chinese AI company’s model, sitting among a cluster of American flagships. Two years ago, nobody would have predicted that position.

There’s another detail on that leaderboard worth lingering on: the cost-per-task column. Opus 5 (High) costs $2.50, Fable 5 (High) costs $2.36, GPT-5.6 Sol (xHigh) costs $1.25, and Kimi K3 (Max) costs just $0.79. The score gap between the top entries is nearly imperceptible to an ordinary user, yet the bill can differ by more than 3x. Competing on the same stage, price is the dimension that actually separates them.

So, for the first category of use case — everyday Q&A, writing, translation — a mid-tier model is more than enough, and if budget is tight, Chinese labs’ models are the best value available. This gap isn’t a capability gap. It’s a premium gap.

For the second category — coding and long-document analysis — flagship and reasoning-tuned models genuinely earn their price, but it’s worth understanding exactly where the money goes. Fable 5.1’s cache read cost dropped from $1 to $0.25 per million tokens, and coding happens to be the scenario that consumes the most cache, since a model has to repeatedly re-read the same codebase — this can account for more than half of total consumption in long tasks. Anyone running long tasks should study cache pricing more closely than benchmark scores.

But don’t max everything out reflexively either. As noted above, running at maximum reasoning intensity is actually more expensive, and using a model at Fable 5.1’s tier for small everyday tweaks is both slow and costly — that 49% slowdown in the CodeRabbit data wasn’t free.

For the third category — building automated workflows — look at the Agent leaderboard, not the benchmark leaderboard. Whether a model can autonomously verify its own results and adjust priorities matters far more than how polished a single response looks. The fact that the top of the Agent leaderboard spans three leading labs plus a Chinese AI company shows the agent space hasn’t consolidated into a monopoly — there’s more room to choose than you might think.

For the fourth category — anything involving sensitive data — check the data retention policy before checking the score. Where your data lives, and whether it gets used for training, matters more and more relative to benchmark performance. Anthropic making data sovereignty a headline feature this round is the industry setting its own direction.

One last piece of advice that isn’t tied to any specific use case: test it yourself. Take a real task from your actual work, run it through two candidate models side by side, and compare the output and the bill. Ten minutes of that tells you more than ten review articles. Marketing language belongs to the vendor. Output and the bill belong to you.

Don’t reverse the order. Define the task first, then choose the model. Don’t go shopping for a problem to hand your most powerful model.

Where the Value Is Headed

A hundred years ago, when electricity first became widespread, the real money wasn’t made by power plants. Power plants eventually became a public utility, with margins as thin as paper. The money was made by the people who used electricity to actually do things.

Large language models are heading down the same road. Once a model is powerful enough and cheap enough, it stops being a money-printing machine and becomes a utility — water, electricity, gas. Price competition will keep squeezing margins at the model layer. Value won’t disappear — it will migrate to two places. One is scarce supply: data. As models keep getting more capable, they’ll depend more and more on high-quality data beyond the public internet. The other is grounded application: the agent layer. The model is the engine. The application is the car that actually drives on the road.

For ordinary people, this might be the friendliest entry point there’s ever been. The price of model capability has been driven down. What separates people now is who has actually thought through their own task, and who holds the gateway to the data.

Back to that Porsche at the opening. A flagship starting to talk about fuel economy isn’t a sign of weakness — it’s a sign it’s about to go after everyone’s market.

The story of competing on benchmarks is over. The story of competing on data is just getting started.