Two of the three AI tools this column has covered since June got new flagship models within two days of each other. Anthropic released Claude Fable 5.1 on 1 September, describing it and its restricted sibling Mythos 5.1 as its “most advanced models for coding and knowledge work.” OpenAI announced GPT-6 Astra on 3 September, calling it “the world’s most intelligent and aligned model,” and it reached Plus subscribers over the following days. Both companies published benchmark scores that mean nothing to anyone running a business, and both new models now appear in the model picker of the app you already use, with conditions attached that I will come to.
If you have spent the past few months building prompts, templates and small routines around one of these tools, this is the week they may have changed underneath you. That is not a reason to panic and it is not a reason to switch. It is a reason to test, and the test is simpler than it sounds.
Benchmarks are not a purchasing decision The launch pages read like exam results. OpenAI reports that GPT-6 Astra scores 98 per cent on an advanced mathematics test and finishes computer tasks in 47 per cent less time than its predecessor, GPT-5.6 Sol. Anthropic reports that Mythos 5.1, the same model as Fable 5.1 with fewer restrictions, designed protein binders with a hit rate near 50 per cent against a typical 10 to 15 per cent.
Impressive, and irrelevant to whether the model can reconcile a supplier statement without inventing a line. The claim that matters to a finance team is buried further down in both announcements. OpenAI says Astra produces polished documents, spreadsheets and presentations and handles multi-step workflows; Anthropic says Fable 5.1 is built for knowledge work and long, multi-step tasks.
The vendors have tested it on their material. The reason is that every business has its own dialect. Your supplier names, your chart of accounts, your habit of writing “GCT incl.” in a description field, the way your bank statement exports with the date in the wrong column.
A model that is brilliant on a public benchmark can still be wrong on the fourth line of your aged receivables. The only benchmark that counts is the one you build from your own files. Same prompt, same files, two models.
Summary from source