Practical guide · Models

Migrate to Haiku 5.5: request changes, response handling, and a runnable example

Move from Haiku 4.5 to 5.5 with adaptive thinking, effort, token recounting, robust text parsing, and staging checks for tools and refusals.

Changing the model name is only the first step in a Haiku 5.5 migration. A client that still sends a fixed thinking budget or assumes the first response block is visible text can fail before the cheaper model has a chance to help.

Treat this as a request-and-response migration, then an evaluation of workload behavior. The example below uses the direct Claude Messages API through Python's standard library so its payload and error handling are visible. It is a runnable integration example, not a claim that LLM Scorebook tested the model with a production API key.

At a glance

  • Replace fixed thinking budgets with adaptive thinking and an explicit effort setting.
  • Recount tokens with the new model and handle responses by block type and stop reason.
  • Test ordinary replies, tools, resumed conversations, refusals, and truncations before changing all traffic.

Change the request deliberately

Use claude-haiku-5-5 on the direct Claude API. The official migration guide identifies it as a fixed model ID with no date suffix or separate alias. Bedrock uses anthropic.claude-haiku-5-5; configure the platform's supported region or inference route rather than copying the direct API string into every client.

The old thinking: {"type":"enabled","budget_tokens":N} configuration returns a 400 on Haiku 5.5. Use {"type":"adaptive"} and output_config.effort. Begin with medium and evaluate alternatives. Changing effort can change token use and behavior, so record it beside the model identifier in every result.

Remove sampling controls such as temperature, top_p, and top_k from the migration request. The official guide describes restricted accepted defaults; omitting the controls is clearer than relying on old configurations to happen to match them. Assistant-message prefill also needs replacement with supported output-format or tool patterns.

Source: official migration guide.

A minimal direct API payload

json
{
  "model": "claude-haiku-5-5",
  "max_tokens": 16000,
  "thinking": {"type": "adaptive"},
  "output_config": {"effort": "medium"},
  "messages": [{"role": "user", "content": "Explain why token price differs from cost per accepted result."}]
}

This request exposes the two important decisions: adaptive thinking and a pinned effort. The output budget is an example, not a recommended allocation for all jobs. Thinking contributes to the budget, so limits tuned for short visible replies on the old model may cause truncation. Re-evaluate budget and latency together.

The companion examples/haiku_request.py reads ANTHROPIC_API_KEY from the environment, sends a Messages request, and prints visible text plus usage and stop reason. Run it with --dry-run first to inspect the payload without an API request or charge. Once the environment variable is set in your own shell, run:

shell
python examples/haiku_request.py --dry-run
python examples/haiku_request.py --prompt "Summarize the difference between tariff and task cost."

A live run uses your API account and incurs charges. Keep secrets out of source files and shared terminal captures. The program neither prints nor stores the key.

Source: migration payload. Example implementation by LLM Scorebook.

Parse blocks and stop reasons

Do not assume content[0].text exists. Iterate through content blocks and select those whose type is text for display. Preserve the complete response for a continuing conversation, including thinking and tool-use blocks. A UI's visible text and the API's complete conversation state serve different purposes.

Thinking blocks are empty by default unless summarized display is requested. The migration guide also says these blocks are bound to the producing account or linked accounts and that modifying earlier system, tool, or message content before replay can invalidate them. Test your actual replay and tenant boundaries. A middleware that “cleans up” history can break an otherwise correct tool loop.

Handle stop_reason: "refusal" explicitly. Treat max_tokens as incomplete output; it should not automatically pass an acceptance test. An empty visible reply likewise needs a defined outcome. The sample program exits with an error for those cases and prints usage so they remain visible in your evaluation records. It does not retry automatically, preventing hidden duplicate attempts and charges.

Source: response and conversation migration guidance.

Recount tokens before trusting the old budget

The official guide estimates roughly 30% more tokens for the same text than Haiku 4.5, with content-dependent variation. Use the new model for counting representative prompts. Include the system instructions, tools, retrieved context, images, and previous turns that your production request actually sends. Counting only the user's latest question misses most of a long-running agent's context.

Revisit three limits separately: the application's own prompt budget, the model's context capacity, and the 100,000-token price boundary. They answer different questions. A request can fit in the context window while entering a more expensive pricing band, and a request with a large output allowance can still terminate before producing useful visible text.

For cost logs, distinguish ordinary input from cache creation and cache reads. Reconcile categories with the provider's usage fields. Do not add cache tokens to a field that already includes them, and do not omit cached content from total prompt length simply because it is cheaper to process.

Sources: token recounting, usage categories.

Test tools as their own migration path

A working text request does not establish a working agent. Test one tool call, the corresponding tool result, and the next assistant response. Preserve tool-use identifiers and complete assistant content. Validate tool arguments before execution, and distinguish model mistakes from a tool's timeout or permission failure.

Computer-use integrations have a separate change: the guide replaces the old computer tool with computer_toolset_20260801 on the documented routes and introduces browser_toolset_20260801 for browser tasks. Do not treat the type name as a complete working agent. Platform availability, SDK beta support, screenshots, execution, and resumed state need their own integration tests. The minimal example in this package intentionally covers text only.

Keep user corrections in user messages rather than hiding them inside tool results. The model should be able to distinguish retrieved content from an instruction by the actual user. In an extraction or browsing workflow, validate the final fields or page state using an independent rule. Tool execution alone does not prove the user's intended outcome.

Sources: tool migration, Haiku prompting guidance.

A staging checklist that exercises real failures

Freeze a representative set of production requests and include a straightforward reply, a missing-information case, a tool round trip, a resumed conversation, a long prompt, and a deliberately small output budget. Add a request that your product should refuse, plus a legitimate request that resembles it. Use your application's real parser and validation rather than a separate demo interface.

Run the incumbent and Haiku 5.5 with identical task inputs and acceptance rules, recording model, effort, route, usage, elapsed time, and outcome. Repeat tasks where nondeterminism affects the decision. Review the failures before tuning: you may need better retrieval, a tool repair, a larger output budget, or a stronger model rather than a more forceful prompt.

Roll out gradually with a reversible route switch. Maintain quality and latency thresholds, and keep escalations in the total charge. The migration is complete when the application's supported flows work and the new route meets its acceptance requirements. A successful HTTP response is the beginning of that check, not its conclusion.

Documentation checked October 8, 2026. Prepared with AI assistance. Example payload and error handling were checked locally; live API execution remains untested.

RUN THE EXAMPLE

Keep the code close.

Python standard library. See the article for offline checks and live-integration limits.

Download haiku_request.py
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology