cd /news/ai-tools/claude-code-s-keep-coding-instructio… · home › topics › ai-tools › article
[ARTICLE · art-148585] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Claude Code's keep-coding-instructions kept 1,044 tokens in the full prompt, 0 in the short one, and 18 of 18 fixes matched

A developer measured Claude Code 2.1.289's `keep-coding-instructions` output-style flag by capturing the API requests it sends, finding the flag controls a 1,044-token "# Doing tasks" section that appears only in the full system prompt and vanishes entirely in the shorter default prompt. Across 40 runs of a one-line bug fix, 18 of 18 runs in each of three styles made the same edit, added no comment and verified the result, with the custom style's main effect being a doubled closing message.

by read19 min views7 publishedOct 10, 2026

On Claude Code 2.1.289 with Opus 5.5, keep-coding-instructions: true decided whether a 1,044-token # Doing tasks section stayed in the system prompt, but only when Claude Code sent its full prompt; in the shorter prompt this account got by default, a custom style with the key and the same style without it produced identical requests. On a one-line bug fix, 18 of 18 runs across three styles made the same edit, added no comment and checked the result after editing, and the custom style's main effect was to double the length of the closing message.

Output styles are how you change the way Claude talks for a whole Claude Code session: shorter answers, explanations next to each change, or a different role altogether. The output-styles page attaches one warning to the custom ones you write yourself. Unless the style file sets keep-coding-instructions: true in its frontmatter, Claude Code leaves out its built-in software engineering instructions, "such as how to scope changes, write comments, and verify work."

That sentence raises two questions for anyone who writes a style and then uses it in a repository. How much text actually disappears, measured in tokens? And does a session without it code differently: skip the tests, clean up code nobody asked about, leave comments everywhere? We wanted both answers from the requests Claude Code sends to the API rather than from what the model says about its own instructions, so we built a small lab and ran it 40 times under claude -p.

Everything below ran on 2026-10-05 between 16:23 and 16:33 UTC with Claude Code 2.1.289 (claude --version) and claude-opus-5-5 (selected with --model opus) on macOS. The output-styles page was fetched with trafilatura at 16:20 UTC the same day and fetched again at 16:34 UTC with no change; the other documentation quotes come from fetches between 16:20 and 16:35 UTC.

These are the sentences the lab tests, quoted as fetched from https://code.claude.com/docs/en/output-styles:

keep-coding-instructions is set to true." keep-coding-instructions: true if you're changing how Claude communicates but still want it coding the same way. Leave it out if Claude won't be doing software engineering."false" The key is not new. The changelog entry for 2.0.37 reads "Output Styles: Added keep-coding-instructions option to frontmatter". What the page does not give is a size for "the built-in software engineering instructions" or a list of them beyond those three examples, and the default is false, so every custom style that omits the line is in the dropped group.

There were three conditions, and every run got a fresh directory copied from the same template:

The style is deliberately about writing, the case the docs describe as "Claude won't be doing software engineering":

---
name: Plain Writer
description: Plain, friendly prose for documentation and explanations
---

You are a writing assistant. Answer in plain, friendly prose paragraphs. Avoid bullet lists and headings unless the user asks for them. Keep sentences short, and define any technical term the first time you use it. End every response with a one-sentence summary that starts with "In short:".

Condition (b) inserts keep-coding-instructions: true after the description line, and nothing else changes. Both versions live at .claude/output-styles/plain-writer.md in their own run directory, so the style name and body are byte-identical and only the frontmatter differs. Each run selects it with a one-line .claude/settings.json, {"outputStyle": "Plain Writer"}. The "In short:" sentence is a marker: if a reply ends with it, the style was in force.

The template is a tiny Node project. sum.js holds three functions. sum has an off-by-one bug (its loop starts at index 1, so it skips the first number), average and formatTotal both call sum, and formatTotal is written in a different style from the rest of the file (var, no semicolons) as bait for cleanup nobody asked for. test/sum.test.js has four node:test tests, three of which fail because of the bug, and package.json maps npm test to node --test. Each run directory got git init and one commit, so git status started clean and git diff afterwards shows exactly what the session changed.

Each run was one invocation:

claude -p "$PROMPT" --setting-sources project,local --strict-mcp-config --model opus \
  --permission-mode acceptEdits --output-format stream-json --verbose \
  --session-id "$SID" --max-turns 40 < /dev/null

--setting-sources project,local keeps our user settings, and the hooks and plugins they enable, out of the runs. --strict-mcp-config with no config file means no MCP servers. acceptEdits lets the session edit files without a prompt. We also unset the CLAUDE_* environment variables that the Claude Code session we were working from exports, CLAUDE_EFFORT among them, so the children ran at their own default effort, which every request body recorded as medium.

One setup detail is worth passing on. Our first pilot also put Bash allow rules in .claude/settings.json, and Claude Code printed Ignoring 7 permissions.allow entries from .claude/settings.json: this workspace has not been trusted. The run directories are new and were never opened interactively, so they were never trusted. We moved the allow rules (npm test, node, git diff and a few more) to .claude/settings.local.json, which the settings page says works differently while the file is untracked: "Claude Code applies its allow rules without the workspace trust step it requires for the committed file." The outputStyle key in the same untrusted .claude/settings.json was applied regardless: the init event reported "output_style": "Plain Writer" in 28 of 28 runs that had the style.

The session transcript records token usage but not the system prompt, and we did not want to rely on asking the model which instructions it can see. The monitoring page documents a way to get the request itself: OTEL_LOG_RAW_API_BODIES, which will "Emit the full Anthropic Messages API request and response JSON as api_request_body / api_response_body log events", with "file:<dir> for untruncated bodies on disk with a body_ref pointer in the event". With CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_LOGS_EXPORTER=console, OTEL_METRICS_EXPORTER=none and OTEL_LOG_RAW_API_BODIES=file:<dir>, every run left the exact JSON of every request it sent in its own folder, and nothing extra appeared in the stream-json output. Everything below about what a prompt contains comes from those files. Token counts come from the transcript under ~/.claude/projects/: input_tokens + cache_creation_input_tokens + cache_read_input_tokens of the first assistant message, which counts the whole first request whether or not part of it was served from cache.

The pilots already contradicted what we expected. With the trivial prompt, the request from style + key was the same text as the request from style, no key, except for the run's directory name, its commit hash and a per-request ID. Neither had a section about doing software engineering tasks, and neither did the default.

The explanation is on the environment variables page. CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT is described as: "Set to 1 to use a shorter system prompt and abbreviated tool descriptions on any model. Set to 0, false, no, or off to opt out even on models where the experiment or server configuration would otherwise enable it." With the variable unset, the requests from this account were the shorter kind: the tool definitions came to 27,700 characters of JSON, and the default style's system prompt was about 6,900 characters with five headings (# Harness, # Session-specific guidance, # Memory, # Environment, # Context management) and a few unheaded paragraphs. With CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=0, the tool definitions grew to 44,195 characters and the system prompt to about 28,300, gaining # System, # Doing tasks, # Executing actions with care, # Using your tools, # Tone and style and # Text output (does not apply to tool calls).

Since the docs say an experiment or server configuration can switch the short prompt on, which one you get may depend on your account, your model or the week. So we ran every measurement under both. From here on, the short prompt means what this account got with the variable unset, and the full prompt means CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=0.

The trivial prompt was Reply with the single word OK. with --max-turns 1, three runs for each of the six combinations, 18 runs in all. All 18 replied OK; the one-word answer won over the style's "In short:" sentence.

Prompt Style First request, tokens (3 runs) Median vs default
full default 28,137 / 28,139 / 28,139 28,139 —
full style, no key 27,240 / 27,241 / 27,237 27,240 −899
full style + key 28,285 / 28,284 / 28,284 28,284 +145
short default 16,030 / 16,036 / 16,037 16,036 —
short style, no key 16,175 / 16,179 / 16,181 16,179 +143
short style + key 16,178 / 16,181 / 16,183 16,181 +145

Repeats of one condition differed by up to 7 tokens, because the directory name, the commit hash and a request ID all appear in the request. After we replaced those three values with placeholders, the three runs of each combination were byte-identical, which leaves five distinct prompts for six combinations: in the short prompt, style, no key and style + key came out as the same text in 6 of 6 runs.

In the full prompt, the key is worth 1,044 tokens (28,284 against 27,240 at the median; 1,043 to 1,047 when we pair the runs by round). The difference between those two requests is exactly one block, the # Doing tasks section: 3,321 characters, 530 words, 12 bullets. Without the key, a custom style's first request was 899 tokens lighter than no style at all, because the section went out and about 145 tokens of style came in.

The style itself cost 143 to 145 tokens in both prompts, and it arrives in two places. In the top-level system prompt, one sentence changes from "You are an interactive agent that helps users with software engineering tasks." to "You are an interactive agent that helps users according to your "Output Style", which describes how you should respond to user queries." That swap happened with and without the key. The style body arrives under the heading # Output Style: Plain Writer in a system-role message that follows the first user message, next to the environment details, and near the end of the same message comes "Plain Writer output style is active. Remember to follow the specific guidelines for this style."

In the short prompt, the key is worth nothing. There is no # Doing tasks section in any of the three short-prompt conditions, the default included, so the key has nothing to keep.

Since the section is what the key controls, here is what it says. It opens by telling Claude that the user "will primarily request you to perform software engineering tasks" and to read unclear instructions in that light. It asks Claude to defer to the user on whether a task is too large, to prefer editing existing files over creating new ones, and to avoid security holes such as command injection, XSS and SQL injection. Scope gets two bullets, the first of which begins: "Don't add features, refactor, or introduce abstractions beyond what the task requires. A bug fix doesn't need surrounding cleanup; a one-shot operation doesn't need a helper." The second rules out error handling and validation "for scenarios that can't happen". Comments get two bullets as well, starting with "Default to writing no comments. Only add one when the WHY is non-obvious". It rules out backwards-compatibility hacks and says where to send users who want help or to give feedback. And it has one rule about questions rather than tasks: "For exploratory questions ("what could we do about X?", "how should we approach this?", "what do you think?"), respond in 2-3 sentences with a recommendation and the main tradeoff." That bullet ends "Don't implement until the user agrees."

On verification, the section has a single bullet, and it is about interfaces: "For UI or frontend changes, start the dev server and use the feature in a browser before reporting the task as complete." Nothing else in it says to run tests.

Some of the same ground is covered elsewhere, in parts of the prompt the key does not touch. The full prompt's text-output section, present in all three full-prompt conditions, still says "In code: default to writing no comments." and "End-of-turn summary: one or two sentences." In the short prompt, all three conditions carry "Write code that reads like the surrounding code: match its comment density, naming, and idiom." and "Report outcomes faithfully: if tests fail, say so with the output; if a step was skipped, say that; when something is done and verified, state it plainly without hedging."

So of the docs' three examples, scope is the one that really leaves with the section. Comment guidance leaves once and stays once. Verification leaves only for UI work, because that is the only verification the section mentions.

The coding task was the prompt fix the bug in sum.js, three runs for each of the six combinations.

Prompt Style Ran npm test after the edit Ran a node -e check instead Tool calls per run Closing message, words
full default 1 of 3 2 of 3 3, 3, 3 50, 54, 39
full style, no key 3 of 3 0 of 3 3, 3, 3 120, 89, 98
full style + key 2 of 3 1 of 3 3, 3, 3 98, 102, 101
short default 3 of 3 0 of 3 3, 3, 4 68, 51, 57
short style, no key 3 of 3 0 of 3 3, 3, 3 104, 106, 105
short style + key 3 of 3 0 of 3 4, 3, 3 111, 104, 117

What the table cannot show is how alike the runs were. The diff was byte-identical in 18 of 18: one line in sum.js, let i = 1 became let i = 0. No run added a comment, no run touched formatTotal, its var or its missing semicolons, no run created a file, and the test suite passed afterwards in all 18 directories (we ran it in each one once the sessions were done). Every run followed the same three steps: look at sum.js (with ls and cat, or the Read tool), make one edit, then check. The two runs with four tool calls split the check into two Bash calls. No run ran the tests before editing: 0 of 18.

Checking after the edit happened in 18 of 18 runs. In 15 it was the project's test suite. In the other 3, Claude imported the module and printed a few results, for instance node -e 'const s=require("./sum.js");console.log(s.sum([1,2,3]),s.average([2,4]),s.formatTotal([5]))', and the closing message said "a quick check" rather than claiming the tests passed. All three of those runs came from conditions that had the # Doing tasks section: two from the full-prompt default, one from full prompt with the key. Three runs per cell is too few to call that a trend, and we don't. It is still the opposite of "drop the coding instructions and the verification goes with them".

The style did show, in the closing message. All 12 styled runs ended with "In short:", none of the 6 default runs did, and the styled closings ran 89 to 120 words against 39 to 68 for the default. They also did what the style asks, defining terms along the way ("An index is an item's position in a list, and JavaScript lists start at 0"). In the full prompt the "End-of-turn summary: one or two sentences." line was present in all three conditions, and the style's instructions won over it in 6 of 6 styled runs. The runs cost $0.109 to $0.117 each with the full prompt and $0.075 to $0.083 with the short one.

The bug fix exercised the scope and comment bullets and found nothing to separate. The dropped section has one rule with a target that can be counted, the one for exploratory questions: two or three sentences, a recommendation, the main tradeoff, no implementation. So we spent the last 4 runs of our 40-run budget on the question what could we do about the formatTotal function in sum.js?, in the full prompt, style without the key against style with it, two runs each.

Style Words Sentences before "In short:" Paragraphs Tool calls Files changed
style, no key 245, 252 16, 17 5, 6 2, 3 0, 0
style + key 174, 140 10, 8 3, 2 1, 1 0, 0

Word counts include the "In short:" line. Neither condition came near two or three sentences, which is not surprising when the same request also tells Claude to write friendly paragraphs and define every term. None of the four replies changed a file, and all four found the real problem (the bug in sum) and offered to fix it. With the section present, though, the replies were 37% shorter on average (157 words against 248.5), each made one tool call instead of two or three, and one of them used the instruction's own word: "The tradeoff is that fixing sum changes what average returns as well." One of the runs without the key went as far as running the test suite to answer a question that asked for an opinion. Two runs per cell make this a hint, not a measurement we would build on, but it is the one place in the 22 behavior runs where the runs with and without the section did not overlap at all.

Add keep-coding-instructions: true to any style you use in a repository where Claude edits code. In these runs it cost 0 extra tokens when Claude Code sent the short prompt and 1,044 on the first request when it sent the full one, and after the first request that part of the prompt is a candidate for the cache like the rest. What it brings back in the full prompt is the scope rule, the comment rule, the security bullet and the rule for answering exploratory questions briefly.

Don't count on the key for verification. In this version, the section it controls mentions verification only for UI and frontend work, and on a non-UI fix all 18 runs checked their work after the edit, with or without it. If a check has to run every time, the output-styles page's own comparison table points elsewhere: for "Something to happen every time without exception" it recommends a hook, because "Claude Code runs a hook itself at a lifecycle event, so it doesn't depend on Claude following an instruction".

Don't leave the key out to save tokens, either. A writing-only style without it was 899 tokens lighter than the default in the full prompt, 3.2% of a 28,139-token first request, and it saved nothing in the short prompt.

Expect the style's own text to change a coding session more than the key does. Here it doubled the closing message, and in the full prompt it overrode a one-or-two-sentence rule that sat in the same request. If a style is meant for documentation work and will also be used while fixing code, say in the style what you want the coding replies to look like.

And look at your own request before trusting any of the numbers above, including ours. What the key keeps depends on which system prompt Claude Code chooses, and the docs say an experiment or server configuration can make that choice. OTEL_LOG_RAW_API_BODIES=file:<dir> and a grep for # Doing tasks answer the question for your account in one run.

Every run was headless, on one model, one Claude Code version and macOS. Other models may get the full prompt by default; we only reached it through the environment variable. We did not test interactive sessions, switching styles mid-session with /output-style, user-level or plugin styles, force-for-plugin, or the built-in styles, which the docs say keep the coding instructions. The task was small: one obvious bug, a test suite that was easy to find, no real pressure to widen the change. A larger repository with unclear scope might separate the conditions where this one did not, and the UI verification bullet went untested because nothing here had a UI. Three runs per cell for the bug fix and two for the question are counts, not rates. We measured only the first request of each session, not what caching does to later ones. We did not ask the model to report its own instructions, since the request bodies made that unnecessary.

The style file and the template project are as described above. One run of the token measurement, from inside a run directory:

printf '{\n  "outputStyle": "Plain Writer"\n}\n' > .claude/settings.json
CLAUDE_CODE_ENABLE_TELEMETRY=1 OTEL_LOGS_EXPORTER=console OTEL_METRICS_EXPORTER=none \
OTEL_LOG_RAW_API_BODIES="file:$PWD/../bodies/run1" CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=0 \
claude -p "Reply with the single word OK." --setting-sources project,local --strict-mcp-config \
  --model opus --permission-mode acceptEdits --output-format stream-json --verbose \
  --max-turns 1 --session-id "$(uuidgen | tr A-Z a-z)" < /dev/null > out.jsonl

grep -c '# Doing tasks' ../bodies/run1/*.request.json

Leave out CLAUDE_CODE_SIMPLE_SYSTEM_PROMPT=0 to see the prompt your account gets by default, and leave out the printf line for the default style. For the token count, open ~/.claude/projects/<project>/<session-id>.jsonl, take the first assistant record's message.usage, and add input_tokens, cache_creation_input_tokens and cache_read_input_tokens. For the behavior runs, swap the prompt for fix the bug in sum.js, raise --max-turns, put the Bash allow rules in .claude/settings.local.json, and read the tool calls from the same transcript and the change from git diff.

Forty claude -p runs on 2026-10-05 between 16:23 and 16:33 UTC, Claude Code 2.1.289 on macOS, claude-opus-5-5, effort medium as recorded in the request bodies: 18 token runs (3 for each of 6 combinations), 18 bug-fix runs (3 for each of 6 combinations) and 4 question runs (2 for each of 2 combinations). All 40 exited 0 with "subtype": "success". The token runs took 1.2 to 3.5 seconds each, the bug fixes 11.2 to 22.5 seconds and the questions 15.0 to 22.7 seconds. The reported cost for all forty was $3.62.

Whether a style keeps Claude Code's coding instructions is one line of frontmatter, and these runs say it is worth checking which system prompt that line is acting on before deciding what it costs.

If your account gets the full prompt by default, or your custom style drops more than 1,044 tokens, post your Claude Code version and the number in the comments below.

── more in #ai-tools 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-s-keep-c…] indexed:0 read:19min 2026-10-10 · —