Claude Opus 5.5 Removes the Off Switch for Thinking
Anthropic's new prompting guide confirms Opus 5.5 rejects thinking: disabled. Effort is now the only dial, and the default moved down from high to medium.
The setting that no longer exists
Buried in Anthropic's prompting guide for Claude Opus 5.5 is a change that breaks working code rather than merely improving on it: the model does not accept thinking: {"type": "disabled"}. Claude Opus 5 allowed it at high effort or below. Opus 5.5 does not. Thinking is always on, and the only remaining control over how much of it happens is the effort level.
For most teams the headline numbers are friendlier. Anthropic says Opus 5.5 produces output tokens more than 30 percent faster than Opus 5 and tends to finish the same task using fewer of them. The company also says existing Opus 5 prompts should work without changes. The catch is in the settings around the prompt, not the prompt itself.
Why "medium" is not the same medium
Effort is the dial that decides how long the model deliberates before answering. Opus 5.5 defaults to medium; Opus 5 defaulted to high. The names do not transfer between models. According to Anthropic's own testing, Opus 5.5 at medium matches or exceeds Opus 5 at high on coding and knowledge-work evaluations, and on several coding evaluations low comes close at much lower cost. That is a vendor benchmark, not an independent one, and the guide's advice is to test levels against your own evals rather than carry over the value you used before.
Copy the old setting across and the bill moves the wrong way. At any given level, the guide says, Opus 5.5 thinks more per turn than Opus 5 did, especially at xhigh and max. Thinking tokens count toward max_tokens even when the thinking content is never returned to you, so a ceiling sized for Opus 5 with thinking switched off can truncate replies. Anthropic suggests 128,000, the model's maximum, for the long turns agentic coding produces, and says lowering the effort level cuts thinking more reliably than instructions telling the model to think less.
One operational trap: changing the top-level effort value between requests invalidates the prompt cache, the stored copy of your repeated system prompt that keeps costs down. A per-message effort change, currently in beta, preserves it.
The agent that stops because it was being polite
The most interesting failure mode in the guide is behavioural. On long multipart tasks, Opus 5.5 keeps the user posted as it works, and some of those updates end the turn with text instead of a tool call, returning stop_reason: "end_turn". An unattended agent loop reads that as "task finished" and shuts down. The work was not finished. The model was reporting.
Anthropic's fixes are mundane and worth copying: keep the task's parts in a checklist the model updates, treat a text-only ending as a report rather than proof of completion, and send a short message naming the open items. Cap automatic continuations at two or three so a genuinely stuck run ends and gets reviewed. If a background command or subagent is still running, wait for its output before deciding anything is done.
The guide also offers a long system-prompt paragraph that names four specific ways the model ends turns prematurely, including summaries that announce the next step without taking it, and offers to continue that wait for an answer nobody will give. Anthropic is explicit that it is for fully unattended agents only, that you should keep your own confirmation step for irreversible actions, and that it costs somewhat more tool calls and output tokens per task. Adding it partway through a session invalidates the conversation's earlier thinking blocks, so it goes in from the first request.
Questions You Should Be Asking
- If thinking can no longer be disabled, what is our actual floor on time to first token for chat-facing features, measured on our own traffic rather than Anthropic's evals?
- Which of our prompts instruct the model to write out its reasoning in the response as a substitute for thinking? Those can now be declined under the reasoning_extraction refusal category.
- Does our agent harness inspect content block types, or does it assume the first block is text? A response may begin with a thinking block whose field is empty under the default display setting.
- How many of our "successful" unattended runs last quarter actually stopped at a progress report we counted as completion?
- Who owns the effort setting, and are we paying for xhigh on tasks where we never measured a quality gain?
What To Watch Next
Watch your token bills at constant effort. If costs rise after migrating despite the claimed 30 percent speed gain and lower token use per task, the cause is almost certainly a carried-over effort value combined with a max_tokens ceiling set when thinking was optional. That single comparison, same workload before and after, will tell you whether Opus 5.5 is cheaper for you or just faster for Anthropic's benchmarks.
- 1Grep your codebase for thinking: {"type": "disabled"} before upgrading to Opus 5.5, since that parameter now throws instead of being ignored.
- 2Replace disabled-thinking calls with the lowest effort level, then re-benchmark latency and cost because minimum thinking still consumes tokens.
- 3Re-test any prompt tuned at 'medium' effort on Opus 5.5, as the effort levels are recalibrated and won't match Opus 5 behavior.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
