Stop Saying 'Think Step by Step': What Addy Osmani's Opus 5.5 Prompting Guide Actually Says Addy Osmani published a guide on Anthropic's engineering blog, "Getting the most out of Opus 5.5 in Claude and Claude Code," arguing that Opus 5.5 now decides how much to think on its own, making older prompt tricks like "think step by step" dead weight. The guide recommends giving the full task in one message with an explicit finish line and stop condition, banning specific unwanted design styles, using CLAUDE.md as run control, fanning out work to subagents with evidence verification, and keeping a task checklist in a file such as TASKS.md. Anthropic's own chat-product testing found that removing a "think carefully" line made replies start sooner with no clear quality drop. Open your saved prompts right now and search for the words "think carefully" or "step by step." If you find them, you are carrying habits from an older generation of models, and according to Anthropic's newest official guide, those habits are now making your replies slower for no gain. Last week Addy Osmani published "Getting the most out of Opus 5.5 in Claude and Claude Code" on the Anthropic engineering blog, and it made the Hacker News front page with real discussion. The theme of the whole guide is uncomfortable for anyone who spent 2025 collecting prompt tricks: the model now decides how much to think on its own, so your job is no longer to push it toward reasoning. Your job is to define "done," tell it when to stop and ask, and then get out of its way. One disclosure before the breakdown: everything below is built from Osmani's guide and Anthropic's own documentation, all linked inline. I have been running my own multi-week experiments in live repos, and those confirmed the guide's core claim before I read it, but the numbers and specific recommendations here are the guide's, not mine. Delete "think carefully" lines. In Anthropic's own testing in a chat product, removing a "think carefully" line made replies start sooner with no clear drop in quality. Opus 5.5 thinks before every reply and decides how much thinking the task deserves. The instruction is now dead weight, and in saved instructions it applies to every single message you send. Give the whole task in one message. This is the biggest change from how most of us work. Osmani's example prompt has three parts, and every one of them matters: Why the change? Early testers ran Opus 5.5 on long coding tasks for hours with little oversight, and its biggest gains over Opus 5 are exactly there, carrying a change through a large repository until the tests pass. A finish line tells it when it is done. A stop condition tells it the one case where you want to be woken up. Everything else should be a status note, not a question. Name the styles you do not want. For design work, "avoid a generic look" mostly swaps one generic look for another. Osmani's example is a list of specific habits to ban: no cream or off-white background, no italic accent words in headings, no numbered "01 / 02 / 03" section labels, no monospace labels, no pill-shaped buttons. Specific negatives beat vague positives, and if you dislike what it picks next, add that to the ban list and rerun. The second section of the guide is the most valuable part for anyone running agent sessions, because it turns CLAUDE.md from a style guide into run control. Osmani's suggested rule is short enough to paste: Two details worth pausing on. First, the guide warns that Opus 5.5 sometimes stops mid-task to report instead of continuing, giving you a summary that names the next step without taking it, or an offer to continue. If your runs keep ending with "Want me to continue?", the rule above is the fix, and the fix for a single occurrence is to just reply "continue." Second, the destructive-command carve-out matters: a keep-going rule means fewer stops, so you keep your own checkpoint before anything risky or hard to undo, and you leave permission prompts on for destructive commands. Subagents for audits and migrations. For work spread across a large codebase, Osmani suggests asking the model to fan out: "Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it." Note that second half. Anthropic is not telling you to trust the parallel reports; they are telling you to have the model itself verify each subagent's evidence before it accepts it, then finish with a single table: service, affected yes or no, and the evidence. Keep the task list in a file. Long runs fill the context window, and Claude Code then summarizes older turns. A checklist in TASKS.md survives that summarization and shows you at a glance what is done and what is left. Read the file, not the scrollback. When a long run ends, Osmani says to look first for anything Claude is waiting on you for: a decision it left open or a change it wants you to approve. Only then read the rest of the summary. His trick for making this repeatable is to standardize the ending in CLAUDE.md: "End every run with three headings: Blocked on me, Changed, Found." Three more checking habits from the guide: Here is the section of the guide that deserves its own post. Opus 5.5 is the first Opus model to launch with Fable-level bio and cyber safeguards, and in Claude apps and Claude Code, most flagged messages do not get refused. They get silently routed to an older model, and your work just continues there, usually without you noticing. Finding security vulnerabilities in source code is still allowed, and everyday questions still work. But Anthropic admits these safeguards can flag legitimate work, and the check covers everything in the conversation, including files and search results. That means a flag can come from earlier content in the session, not just your last message, and if you do not know the recovery steps you may be getting Opus 5-quality answers while believing you are on Opus 5.5. The recovery steps from the guide: If you do security review or pen-test adjacent work, set that toggle to ask-first today. Getting silently downgraded mid-session is the kind of failure you can spend hours not noticing. One related prompting note: do not ask the model to reproduce its internal reasoning in the reply. That request is one of the flag categories and can be declined outright. Ask for the useful output instead: "Explain why you chose this approach in three sentences." Back-and-forth work, where you read each reply before sending the next message, is where speed actually matters, and that is exactly what fast mode targets: type /fast in Claude Code. You get the same model with text arriving sooner, but it is a research preview, needs extra usage turned on, and costs more per token than standard mode. For long autonomous runs, the opposite logic applies: the run is not interactive, so pay standard rates and let it cook. Run through this before your next long task. This is the part worth bookmarking. Before you hit send: Before a long Claude Code run: Before you trust the result: Once, today: That last one is the meta-lesson of the whole guide. Prompting advice has a shelf life, and the tricks that made Opus 5 feel controllable are now friction. The winning move in 2026 is fewer words in the prompt and more structure in the config: a defined finish line, a written stop rule, and settings you chose on purpose. I write about AI tooling, agentic workflows, and what actually works in practice every week. Subscribe, it is free, and it keeps the deep dives coming. Have you run Opus 5.5 on a long autonomous task yet? Did it stop too often to ask permission, or did it run clean? I want to hear how the "define done and let go" style is working for other people, because it is a real behavior change from how we all worked last year.