In a restaurant kitchen, everyone always talks about the chef. Rarely about those who peel, chop and prepare behind the scenes: the commis chefs. Without them, though, no plate ever leaves the pass.
In AI, small models are the commis chefs. They sort, summarise, classify, by the million, while the big models do the chef's work. Little is said about them. Yet they keep a good share of the services you use running.
On Wednesday, Anthropic released its own: Claude Haiku 5.5.
The price, first
It is the third model in the new Claude 5.5 family in a month, after Opus 5.5 on 22 September and Sonnet 5.5 on the 28th. And it is its price that stands out most.
For a request of under 100,000 tokens, Haiku 5.5 costs 0.10 dollars per million tokens on input and 0.50 on output. Its predecessor, Haiku 4.5, cost 1 and 5 dollars. Ten times cheaper, then.
Above 100,000 tokens, the rate rises to 0.50 and 2.50 dollars, still half what it was before. The reason for that threshold is simple: a very long request requires far more computation, like reading an entire book rather than a single page. According to Anthropic, around 90% of requests sent to Haiku 4.5 stayed below that threshold, and the company estimates that a typical workload will cost around 75% less.
What it is worth
Anthropic presents it as its fastest and most capable small model to date. According to the company, it outperforms Haiku 4.5 and GPT-6 Luna, its direct competitor at OpenAI, on the tests it selected. Sonnet 5.5, the mid-range model, remains ahead.
Companies that tested it before release report clear gains. One of them reports a score 11 points higher than Haiku 4.5, with responses twice as fast. As always, these figures come from the vendor and its partners: independent testing will be needed to confirm them.
A first for a model of this size: its effort level can be adjusted, meaning the time it takes to think before answering. More effort for a delicate task, less for a simple question.
Compare the rate cards. Claude Haiku 5.5 and GPT-6 Luna, the small models from Anthropic and OpenAI: 0.10 dollars on input and 0.50 on output, each. Claude Sonnet 5.5 and GPT-6 Sol, their mid-range models: 2 dollars and 10 dollars, each. The two great AI rivals now display exactly the same prices, tier by tier. Fierce competition that ends up looking alike, like two petrol stations facing each other on the same road.
What it is for
Haiku 5.5 targets repetitive, high-volume work: summarising documents, classifying information, querying databases. And above all, acting as a "sub-agent": a large model hands a small task to a small one, faster and cheaper.
One Anthropic customer gives a telling example: while a large model builds a presentation, a Haiku sub-agent goes into a company's annual report to find the precise figure the presentation needs. The chef cooks, the commis chef fetches the ingredients.
The model is available to developers, and in the Claude app for all users, including those on the free version.
What we take away
For most people, this change will be invisible. But it matters. When the clerk becomes ten times cheaper, every service built on top of it can become faster and cheaper.
It also confirms a trend we have been following since September: prices are collapsing and versions keep coming. With a paradox we detail this afternoon: never have small models been so good and so cheap, and yet Google is cutting its free version back to its smallest model.