Abstract:
Anthropic officially releases Claude Haiku 5.5, which is the third model in the Claude 5.5 series and the new generation of Haiku model launched nearly a year after Claude Haiku 4.5. Anthropic positions it as a lightweight model for high-concurrency, low-latency and cost-sensitive tasks, and by significantly reducing the API price, Haiku 5.5 directly enters the competitive ranks of current low-cost AI models.

According to the prices announced by Anthropic, for requests within 100,000 Tokens, the input price of Claude Haiku 5.5 is US$0.10 per million Tokens, and the output price is US$0.50 per million Tokens. Compared with the prices of the previous generation Haiku 4.5, which were US$1 and US$5 respectively, the book price has been directly reduced by 90%.
If the request exceeds 100,000 Tokens, the price of Haiku 5.5 will increase significantly, with input and output prices rising to US$0.50 and US$2.50 per million Tokens respectively. In other words, once the threshold of 100,000 Tokens is exceeded, the input and output prices will become 5 times the basic price.
Anthropic stated that based on actual workload calculations, the average operating cost of Haiku 5.5 is approximately 75% lower than that of Haiku 4.5. One of the important reasons why the average drop has not reached 90% of the advertised price is that Haiku 5.5 uses a new word segmenter.
The new tokenizer is similar to the tokenization method used by Claude Sonnet 5.5 and Opus 5.5. The same piece of text may be calculated as more Tokens in Haiku 5.5. Anthropic's own instructions show that the new word segmentation method will slightly increase the number of tokens generated when completing the same task; some actual tests have found that the same long text may consume approximately 25% to 30% more tokens in Haiku 5.5 than in Haiku 4.5.
This means that Haiku 5.5 has a "hidden cost" that is easily overlooked. Although the price per million tokens has dropped significantly, the actual fee users pay also depends on how many tokens the model consumes to complete the task. Therefore, simply comparing the price per million Tokens cannot fully reflect the actual operating costs of the two generations of models.
Haiku 5.5 still has a very large context window, supporting up to 1 million Tokens, and the maximum output length reaches 128K Tokens. It also supports adaptive thinking capabilities that can dynamically decide how much reasoning resources need to be invested based on the complexity of the task.
Anthropic specifically emphasized that the focus of Haiku 5.5 is not to compete with Opus 5.5 or Sonnet 5.5 for all complex tasks, but to undertake a large number of high-frequency, repetitive and response-speed-sensitive work. For example, text summarization, data classification, information extraction, database query, request routing, context compression, etc. are all suitable application scenarios for Haiku 5.5.
In the field of AI Agent, Haiku 5.5 is also designed as an auxiliary model for large models. Anthropic recommends that developers can let Sonnet 5.5 or Opus 5.5 be responsible for complex main tasks, and then let Haiku 5.5 serve as a sub-agent to handle a large number of simple repetitive operations, thereby reducing the operating cost of the entire Agent system.

Code development is also one of the important application directions of Haiku 5.5. Anthropic said that the model can be used as a coding sub-agent of the large Claude model to handle code searches, simple modifications, and other repetitive tasks that can be split out. However, in complex multi-step Agent programming tasks, Sonnet 5.5 and Opus 5.5 still have more obvious advantages.
Speed is another big selling point of Haiku 5.5. Anthropic says this is the company's most responsive model yet, making it particularly suitable for real-time customer service, browser operations, and applications that require quick feedback.
In addition to lowering the price of Haiku 5.5 itself, Anthropic also simultaneously adjusted the cache read price of Claude Sonnet 5.5. The cache read fee is reduced from US$0.20 to US$0.10 per million Tokens, a decrease of 50%. Anthropic believes that this change can reduce the actual cost of Sonnet 5.5 in most agent workloads by about 20%, since agent tasks often repeatedly read cache contents.
What’s most interesting about the Haiku 5.5, though, is its competition with lower-priced models. Based on the standard price of less than 100,000 Tokens, Haiku 5.5 is already in the same price range as OpenAI’s latest low-cost model. This means that developers can now obtain a model with a context window of millions of Tokens, support for adaptive reasoning and strong coding capabilities at a very low token price.
On the other hand, the price dividing line of 100,000 Tokens may also become an issue that developers need to pay special attention to. For ordinary short text requests, this threshold usually has no impact; but for scenarios involving long document processing, complex Agents, multiple rounds of code tasks, and a large amount of context input at the same time, the model may easily approach or even exceed 100,000 Tokens. Once this limit is exceeded, both input and output prices will directly increase by 5 times.
This also makes Token efficiency itself become more and more important. As the price of AI models continues to decline, developers who used to only need to focus on “how much does it cost per million tokens” must now also consider how many tokens the model consumes to complete the same task. A model that has a lower unit price but requires generating more Tokens to complete the task may not always have a lower actual final cost.
Anthropic currently provides Haiku 5.5 on the Claude platform and channels such as Amazon Web Services, Google Cloud, and Microsoft Foundry, and has also added it to Claude Code. Ordinary users can use Haiku 5.5 on Claude's web page, iOS and Android apps.
Looking at the entire Claude 5.5 product line, Anthropic has formed a relatively clear division of labor: Opus 5.5 is responsible for the most complex reasoning and agent tasks, Sonnet 5.5 is responsible for the balance between performance and cost, and Haiku 5.5 is further focused on high concurrency, low latency and low-cost scenarios.
Therefore, the real competitiveness of Haiku 5.5 is not just "cheapness", but Anthropic's attempt to use a lower unit price to push a model with strong reasoning and agent capabilities into large-scale calling scenarios. However, the new token measurement method and the price jump after 100,000 tokens also mean that enterprises must calculate the total cost based on the actual workload before deployment, rather than just looking at the advertised price per million tokens.
If AI Agents are further popularized in the future, a large number of model calls will change from one question and answer today to dozens or even hundreds of continuous calls. In this case, the price of a single call, response speed and Token consumption efficiency will directly affect the operating costs of the entire AI system. Haiku 5.5 was launched under this background. Its real goal is probably not to become the most powerful Claude model, but to become the "low-cost engine" that bears the largest call volume in Anthropic's entire AI ecosystem.
Comments