The documentation for Claude Opus 5.5 explains that the model generates output tokens more than 30 percent faster than its predecessor and tends to finish the same task with fewer tokens, according to the guide. Existing Claude Opus 5 prompts should work without changes, but the guide recommends starting with the medium effort level, which is the default for Opus 5.5, whereas Opus 5 defaults to high. In Anthropic’s testing, medium effort on Opus 5.5 matches or exceeds high effort on Opus 5 for coding and knowledge‑work evaluations, and low effort can come close to that performance at much lower cost. Because thinking is always on, effort is the primary knob for trading intelligence, latency and cost; users should set max_tokens high enough to accommodate thinking tokens, with 128,000 tokens working well for long agentic coding turns. If the current Opus 5 integration ran with thinking disabled, four changes are advised: start at low effort and measure, remove any prompts that asked the model to write out its reasoning as a substitute for thinking, re‑test any thinking‑disabled mitigations, and read the response by block type rather than assuming the first block is text. For unattended agentic runs that stop partway through a long task after reporting progress, the guide suggests adding user‑facing progress updates at predictable points. When a safeguard refusal occurs (stop_reason: ‘refusal’), developers should follow the refusal handling patterns outlined in the document. In multi‑app workflows where an agent misses information not explicitly pointed to by the task, exploring context across connected apps is recommended. For teams of agents that should finish sooner, using time signals in the harness can help. If replies in a chat application start slowly because the model thinks at length first, adding a system‑prompt line such as ‘Answer directly without deliberating.’ can reduce thinking, though quality should be measured when the line is added. For complex visual inputs like dense charts, diagrams or screenshots, the model reads visual material more accurately than Opus 5 without extra tooling; even at its lowest effort setting it can read values off dense charts more accurately than Opus 5 did at its highest effort, using a small fraction of output tokens. The guide also notes that frontend output may look generic, and adjusting frontend design defaults can improve results. All recommendations are drawn from the official prompting guide for Claude Opus 5.5.

