TL;DR: Anthropic released claude-sonnet-5-5 on September 28, 2026 at the same $2/$10 pricing as Sonnet 5. The model runs 30% faster, lowers cost per task by up to 30% through fewer tokens, and introduces five breaking API changes. The two most likely to affect compliance workflows are forced tool calls now returning an error and thinking-off requiring a new parameter. The governance upgrade that matters most: categorized refusals with five machine-readable labels, which turns model declines into auditable structured data for the first time.
Anthropic released Claude Sonnet 5.5 on September 28, 2026 with an unchanged rate card and five breaking API changes. For compliance teams that have built document review pipelines, policy Q&A systems, or regulatory obligation extractors on Sonnet 5, the question is not whether to migrate eventually (Anthropic is signaling that direction clearly), but whether to migrate now, and which changes require engineering attention before the switch.
This guide covers the breaking changes and their compliance workflow risk, the governance-relevant additions (particularly categorized refusals and the knowledge cutoff update), and a six-step migration checklist written for regulated deployments rather than general agent development.
Sonnet 5.5 vs Sonnet 5: Spec Sheet
Same price and context window, fresher knowledge, stricter API behavior. From Anthropic's model documentation:
| Spec | claude-sonnet-5-5 | claude-sonnet-5 |
|---|---|---|
| Released | September 28, 2026 | June 30, 2026 |
| Status | Latest | Legacy (Anthropic recommends migrating) |
| Supported until | At least September 28, 2027 | At least June 30, 2027 |
| Context / max output | 1M / 128K tokens | 1M / 128K tokens |
| Knowledge cutoff | June 2026 | January 2026 |
| Input / output price | $2 / $10 per million tokens | $2 / $10 per million tokens |
| Cache write (5 min / 1 hour) | $2.50 / $4 per million | $2.50 / $4 per million |
| Smallest cacheable prompt | 512 tokens | 1,024 tokens |
| Minimum thinking setting | between_tools | disabled |
| Forced tool choice | Returns error | Allowed |
| Categorized refusals | Yes (5 machine-readable labels) | No |
| Per-message effort | Yes | No |
The smaller cache threshold is quietly useful for compliance deployments. System prompts that carry policy frameworks, role definitions, or regulatory context and fall between 512 and 1,023 tokens are now cacheable for the first time. Cache hits run at $0.20 per million versus $2.00 for uncached input: a 90% reduction on that portion of every call.
The 5 Breaking API Changes, Ranked by Compliance Workflow Risk
Anthropic published five breaking changes at launch. They are not equally likely to surface in a compliance deployment:
1. Forced tool choice now returns an error (High risk)
Sonnet 5 accepted tool_choice: {type: "tool", name: "extract_clause"}, which told the model it must call a specific function. Claude Sonnet 5.5 rejects this and returns an error.
What breaks: Document extraction and structured output pipelines that force a specific tool call, common in contract review systems built around extract_parties(), classify_obligation(), or named extraction functions. These fail immediately on Sonnet 5.5 without code changes.
Fix: Replace tool choice with auto plus strict tool definitions. Add output validation in application code to confirm the expected tool was called. This is more robust anyway, since forced tool selection was never a guarantee of correct output.
2. Thinking off requires between_tools instead of disabled (Medium risk)
Sonnet 5 accepted "thinking": {"type": "disabled"} to suppress chain-of-thought reasoning. Sonnet 5.5 only accepts between_tools at Low, Medium, or High effort. Passing disabled at Xhigh or Max returns an error.
What breaks: Deployments that suppress thinking to reduce latency on routine classification tasks. At Xhigh or Max effort, thinking suppression is no longer available at all. Factor that into latency budgets for compliance-critical workflows that previously ran extended reasoning and then discarded it.
Fix: Replace disabled with between_tools. If you were running at Xhigh effort specifically to get higher-quality output with thinking suppressed, reconsider whether High effort meets the quality bar; it typically does for well-defined compliance extraction tasks.
3. Thinking blocks tied to their originating conversation (Medium-low risk)
Thinking blocks generated by Sonnet 5.5 cannot be replayed in a different model, conversation, or account context.
What breaks: If your audit trail stores model reasoning blocks as structured objects and you replay them into a new session (for human review, escalation, or secondary analysis), that pattern breaks. The blocks cannot be reused across session boundaries.
Fix: Log the text content of thinking blocks rather than the thinking block objects themselves. The reasoning is preserved; only the replayable object is restricted.
4. Computer use tool computer_20251124 no longer accepted (Low risk for most)
Teams using computer use on the Claude API or Google Cloud must upgrade to computer_toolset_20260801. Amazon Bedrock users are unaffected.
What breaks: Automated interactions with web-based regulatory portals, form filing systems, or compliance platforms. Most text-based compliance workflows do not use computer use and are unaffected.
Fix: One-line change in tool configuration. No logic changes required.
5. Advisor model restrictions (Low risk for most)
Opus 4.8, Opus 4.7, and Sonnet 5 are no longer accepted as advisor models. Use Sonnet 5.5 or Opus 5.5.
What breaks: Multi-model advisory patterns where a secondary model evaluates or scores primary model output. Relatively uncommon in compliance deployments.
Fix: Update the advisor model designation. No prompt changes required.
Behavioral Changes That Do Not Error Out
Three changes alter model behavior without failing any request:
Effort levels mean something different. Anthropic recalibrated the tiers. Medium on Sonnet 5.5 is not the same computational depth as Medium on Sonnet 5. For routine compliance tasks (document classification, clause tagging, routine Q&A), Medium or Low is now the right setting and costs less. For high-stakes analysis (regulatory obligation mapping, multi-jurisdiction risk scoring), stay at High.
Progress notes move to thinking blocks. Text written between tool calls now comes back in thinking blocks rather than as visible assistant text. A compliance interface that streams model reasoning as it processes a document will go quiet between tool calls until you set a display value or switch to between_tools. Deployments that do not surface intermediate reasoning are unaffected.
Refusals now carry machine-readable labels. This is the governance addition compliance programs will actually care about. Every refusal in Sonnet 5.5 carries one of five structured labels: cyber, bio, frontier_llm, reasoning_extraction, or general_harms. A request that Sonnet 5 declined with unstructured text like "I can't help with that" now comes back with a queryable category.
For compliance programs that log AI interactions, those labels are directly aggregatable. You can answer questions like: "How many requests in the past 30 days were declined for general_harms? Which teams generated them? Is the rate trending up?" That audit trail capability did not exist in Sonnet 5. Add refusal_category to your interaction logging schema before migrating.
Knowledge Cutoff Update: What Sonnet 5.5 Knows That Sonnet 5 Does Not
Sonnet 5's training data runs through January 2026. Sonnet 5.5 extends that to June 2026, adding five additional months of regulatory and legal developments in the model's base knowledge.
Several significant obligations entered into application between January and June 2026. The EU AI Act's prohibited practices provisions became effective February 2, 2026. Sonnet 5 was trained before those provisions applied; Sonnet 5.5 has that enforcement context baked in. A policy Q&A system built on Sonnet 5.5 will reason about EU AI Act prohibitions with a more current baseline.
The first-quarter 2026 wave of state AI enforcement actions, FTC guidance updates, and EU DPA decisions applying the AI Act's transparency obligations also fall in this window. For teams using AI assistants without retrieval augmentation (such as raw chat interfaces rather than RAG-based systems), the fresher cutoff reduces the risk of a legally stale answer.
The practical effect is smaller for well-built compliance systems that supply regulatory text as context. If your system explicitly provides the relevant statute, guidance, or enforcement order in the prompt, the model's training cutoff is secondary. If it does not, Sonnet 5.5's fresher baseline matters.
Nothing after June 2026 is in either model. California Adam's Law (signed September 10, 2026), Colorado's revised ADMT rules (September 23, 2026), and the UN Security Council AI session outcomes all need to be provided explicitly regardless of which Sonnet you run.
Cost for Compliance Workloads
Anthropic kept the $2/$10 rate card unchanged. Actual bills can fall because Sonnet 5.5 typically finishes tasks with fewer tokens and fewer tool-call rounds.
Published estimates put savings at up to 30% per task for agentic workflows. For compliance-specific workloads like contract review, regulatory gap analysis, and policy Q&A, the savings depend on effort level and document complexity. Medium-effort extraction tasks see more savings than High-effort analytical tasks.
The smallest cacheable prompt dropping from 1,024 to 512 tokens is the structural cost change most compliance teams will benefit from. Long system prompts that carry jurisdiction-specific regulatory context, role definitions, and output formatting requirements now hit cache earlier. At $0.20 per million for cache hits versus $2.00 for input, a system prompt called 10,000 times per month turns cache eligibility into real budget savings.
Do not extrapolate from developer benchmarks directly. Published token counts come from agentic coding workflows, not compliance document review. Run your actual workload with both models, compare cost per document reviewed or per obligation extracted, and use that for budget modeling. Track cost per completed task, not cost per token, because effort level changes what counts as "done."
Six-Step Migration Checklist for Compliance Deployments
Swap the model ID to claude-sonnet-5-5, then work through this list:
Step 1: Audit forced tool use. Search your codebase for tool_choice with type tool. List every workflow that forces a specific tool call. Replace with auto and add output validation in application code. Test against your full document corpus, not just representative examples. Edge cases in real contracts surface different tool selection patterns than test documents.
Step 2: Update thinking suppression. Replace "thinking": {"type": "disabled"} with {"type": "between_tools"} wherever you suppress reasoning. Flag any calls running at Xhigh or Max effort with disabled. Those need a policy decision about whether the task should move to High effort or whether Xhigh is genuinely required.
Step 3: Fix thinking block logging. If your audit log captures thinking block objects for replay or review, switch to capturing the text content instead. The reasoning content is identical; only the replayable object is restricted to its originating conversation.
Step 4: Update advisor models. Replace any advisor designation pointing to Opus 4.8, Opus 4.7, or Sonnet 5 with Sonnet 5.5 or Opus 5.5.
Step 5: Recalibrate effort by task type. High effort for multi-jurisdiction regulatory analysis and complex obligation mapping. Medium for well-defined extraction tasks with clear schemas. Low for document routing and classification. Measure cost per task after recalibration. This step often produces the most meaningful cost reduction.
Step 6: Add refusal category logging. Before migrating, update your interaction logging schema to capture refusal_category. This is the highest-value governance addition in Sonnet 5.5 and it costs nothing to capture. Build the logging first, then migrate, so you have structured refusal data from day one on the new model.
Test on a representative sample of your actual compliance workload, not just positive cases. Sonnet 5.5 can read Sonnet 5 thinking blocks, so conversations that migrate mid-session preserve their reasoning history.
When to Stay on Sonnet 5
Sonnet 5 remains supported until at least June 30, 2027. Holding is reasonable if any of these apply to your deployment:
You rely on forced tool calls you cannot rewrite before your next compliance audit cycle. Keeping a working extraction pipeline running while queuing the migration work is a reasonable call.
You need thinking disabled at Xhigh or Max effort specifically. This configuration no longer exists in Sonnet 5.5. If your workflow depends on extended computation with suppressed visible reasoning, that pattern is gone.
Your workflow edits earlier conversation turns rather than appending. Sonnet 5.5 is designed for append-only conversation management; history editing requires staying on Sonnet 5 until that pattern can be refactored.
You use Opus 4.8 or Opus 4.7 as an advisor and cannot update those integrations before your next deployment window.
The real cost of staying is the knowledge cutoff. For compliance use cases that involve EU AI Act obligations that became applicable in February 2026, Sonnet 5's training ends before those provisions went live. A well-built RAG system that supplies relevant regulatory text mitigates this; an unaugmented chat interface does not.
Related Reading
- AI Vendor Due Diligence Checklist , evaluate model providers before committing your compliance workflows
- AI Vendor Contract Red Flags , what to watch for in AI provider agreements when model behavior changes
- AI Spend Governance: Token Budget Controls , manage cost per task across model generations
- Anthropic 19-Day Fable 5 Shutdown: Vendor Dependency Lessons , what Anthropic's export restrictions revealed about model vendor risk
- Does Your AI Vendor Train on Your Data? , Anthropic's training data policy in context
