The 100-metre final. In lane 5, the champion, the one you pay a fortune to bring in. In lane 4, the runner everyone knows, the everyday meet regular.
For years, the gap was visible to the naked eye. Then one night, the two cross the line together, and you have to wait for the photo finish to separate them.
That is roughly what Anthropic has just announced with Claude Sonnet 5.5, released on Monday 28 September.
Six days after its big brother
At Anthropic, models come in families of three sizes. Opus is the top of the range, the most capable and the most expensive. Sonnet is the mid-range, the one most people use day to day. Haiku is the lightest.
The 5.5 family began on 22 September with Claude Opus 5.5, which we analysed that same evening. Sonnet 5.5 arrives six days later. Haiku 5.5 is promised "in the coming weeks".
The price does not move: 2 dollars per million tokens read, 10 dollars per million tokens written, exactly like Sonnet 5. A token is a piece of a word, the unit models are billed in. Opus 5.5 costs twice as much: 4 and 20 dollars.
Anthropic also announces a model more than 30% faster than Sonnet 5, and up to 30% cheaper per task. That last figure does not come from a price cut, but from a model that does the same job in fewer steps.
The 44-profession test
The most telling figure in the announcement comes from a test called GDPval-AA. It does not set maths puzzles: it gives models real work tasks, drawn from 44 professions and nine major sectors. Writing a memo, preparing a spreadsheet, building a presentation.
The result is a points score. Sonnet 5 scored 1,449. Sonnet 5.5 scores 1,844. Opus 5.5, the top model, scores 1,846.
Two points apart. The width of a photo finish.
The same trend shows on a test where the model has to use a computer the way you would, clicking through software: 80.1% success for Sonnet 5.5, 81.8% for Opus 5.5, against 57% for Sonnet 5.
These scores are the ones published by Anthropic. They are no substitute for an independent measurement, and we come back to that below.
What it changes for you
If you use Claude in the app, you will not see a bill change. What you will mostly see is answers arriving faster, and fewer back-and-forths on long tasks. In the apps, the model runs at "medium" effort by default: it does not think for ages when it doesn't need to.
The companies that tried it before release give concrete figures. At Box, which stores documents for millions of businesses, the model was 2.4 times faster and used 12% fewer tokens. At Zendesk, the customer-service specialist, requests were handled 20% faster.
The most striking comes from Base44, a platform that builds apps from a simple description. Across 118 real apps, Sonnet 5.5 needed 3.6 attempts on average to get there. Opus 5, July's top model, needed 7.7.
If you pay for AI by usage, at work or through a tool, this is the kind of figure that matters. The listed price does not tell you what you will pay: what counts is the number of steps to reach the end. It is the same logic that explains why ChatGPT can be free when it costs a fortune.
Sonnet 5.5 is the first Sonnet shipped with filters that prevent its reasoning from being extracted. The technique being targeted is called distillation: you ask a model millions of questions, harvest its detailed answers, and use them to train a competitor on the cheap. It is a bit like filming every one of a champion's training sessions to copy their technique. Anthropic accused Alibaba and DeepSeek of it this year. And on the very day of the release, two Americans clashed in public: for Treasury Secretary Scott Bessent it is "theft"; for Nvidia's boss Jensen Huang, "that's called competition".
Where the champion keeps the edge
Anthropic spells out the limit itself, and credit to it for that: Opus 5.5 remains "clearly stronger" at complex, open-ended work, the kind that requires judgement over time. Sonnet 5.5 is presented as the best choice for "well-scoped" tasks: fixing a bug, producing a polished document, filling in a spreadsheet.
In other words, on a well-marked sprint, Sonnet will do. On a marathon full of surprises, where decisions must be made alone for hours, Opus stays ahead.
One figure also calls for caution. On a command-line programming test, Sonnet 5 scored 10.3% and Sonnet 5.5 reaches 70.6%, more than Opus 5.5. A leap like that, from one version to the next, probably says as much about the test as about the model. We have already seen rankings flip when a third party redoes the measurements.
The coming days will tell whether independent evaluators confirm the two-point gap with Opus.
The field speeding up
Two years ago, choosing an AI model was like choosing a car: the more you paid, the better it was. This release tells a different story.
For the vast majority of what you ask an AI to do, writing, summarising, organising, the mid-range model has caught up with the top of the range. The gap no longer lies in what the model can do, but in the kind of problem you hand it.
The day the everyday runner finishes in the same photo as the champion, it isn't the champion who has slowed down. It's the whole field that has sped up.