# When a Claude Code skill's reference.md and scripts actually load (the 'executed, not loaded' script was read in 13 of 16 runs)

> Source: <https://dev.to/rulestack/13-of-16-claude-code-runs-loaded-the-skill-script-the-docs-call-executed-not-loaded-5cic>
> Published: 2026-10-04 02:17:00+00:00

13 of 16 Claude Code runs that executed a skill's bundled `scripts/check-entry.sh` also loaded its source into context, with `Read` or `cat`, although the skills docs label such a script "executed, not loaded". The rest held as documented on Claude Code 2.1.285 across 20 `claude -p` runs: `reference.md` stayed out of the first request and out of the invocation message, both the same size to the token whether the file held 1,182 or 25,405 characters, and the model read it in 16 of 16 runs whose `SKILL.md` pointed to it and in 0 of 4 whose `SKILL.md` did not.

A skill does not have to be one file. The Claude Code docs suggest keeping `SKILL.md` short and moving the details into a `reference.md`, the examples into their own files and any repeatable logic into `scripts/`. The deal they describe is that the extra files cost nothing until they are needed, and that scripts are run rather than read. Before splitting a skill that way, I wanted to know how literally to take that deal. Three questions decide it. Does a large bundled file cost anything before the model reads it? Will the model find a file that `SKILL.md` does not mention? And does a script's source stay out of the model's context when the model runs it?

So I built one small skill with three supporting files, put a marker string in each file that exists nowhere else, and ran the skill 20 times under `claude -p` with four different ways of pointing at the files from `SKILL.md`. The model's final answer shows some of what it read. The session transcript under `~/.claude/projects/` shows all of it, along with the token count of every request, and every number below comes from the transcripts.

All runs happened on 2026-09-30 with Claude Code 2.1.285 (`claude --version`). Eighteen used `--model opus`, which resolved to `claude-opus-5-5`, and two used `--model sonnet`, which resolved to `claude-sonnet-5-5`. The documentation quotes come from the Claude Code skills page (`https://code.claude.com/docs/en/skills.md`) and from two Agent Skills pages on platform.claude.com, all fetched the same day with `trafilatura`.

The Claude Code skills page covers this in a short section called "Add supporting files":

Skills can include multiple files in their directory. This keeps `SKILL.md` focused on the essentials while letting Claude access detailed reference material only when needed. Large reference docs, API specifications, or example collections don't need to load into context every time the skill runs.

A directory tree follows, with a label on each file. `reference.md` and `examples.md` are "loaded when needed". The script gets a different label:

Then the page says how to wire the files up: "Reference supporting files from `SKILL.md` so Claude knows what each file contains and when to load it". Its example is an "Additional resources" section with two markdown links, `For complete API details, see [reference.md](reference.md)` and `For usage examples, see [examples.md](examples.md)`.

The same page says that Claude Code skills "follow the Agent Skills open standard", and the Agent Skills overview on platform.claude.com is blunter about scripts: "When instructions mention executable scripts, Claude runs them through bash and receives only the output (the script code itself never enters context)." The Agent Skills best-practices page lists "Save tokens (no need to include code in context)" among the benefits of utility scripts, and asks skill authors to "Make clear in your instructions whether Claude should" execute a script, with "Run `analyze_form.py` to extract fields" as the example, or read it as reference.

The skill is called `changelog-entry` and lives in `.claude/skills/changelog-entry/` of a throwaway project. Besides the skill, the project holds a README and a `.claude/settings.json` that sets `disableBundledSkills` to `true`, so the skill listing had exactly one entry. The frontmatter has only a `name` and a one-sentence `description`. The body gives three rules: write one past-tense bullet that starts with an area tag such as `[parser]`, end it with the entry code `(BODY-M3V8)`, and do not create or edit any file.

Each supporting file has its own marker, and using a file changes the answer in a way you can see:

| File | Size | Marker | What it adds to the entry | 
|---|---|---|---|
| `SKILL.md` | 12 to 18 lines | `BODY-M3V8` | the entry code | 
| `reference.md` | 35 lines, 1,182 characters | `REF-Q7X4` | a second code, `(train REF-Q7X4)` , a`BREAKING:` rule and "No period before the codes" | 
| `examples/sample-entry.md` | 8 lines, 261 characters | `EXM-K2P9` | an indented footer line, `Changelog-Set: EXM-K2P9` | 
| `scripts/check-entry.sh` | 19 lines, 595 bytes | `SRC-W5N3` , in a comment | prints a footer line ending in `Checked: RUN-H8T6` | 

`reference.md` says of the second code: "The release script rejects any bullet that does not carry the train tag." The script is what makes the test work. Its source carries the comment `# Source marker: SRC-W5N3`, and it builds its output marker at run time with `printf 'RUN-%s%s' 'H8' 'T6'`, so the string `RUN-H8T6` exists only in its output. When `SRC-W5N3` shows up in a transcript, the script's source reached the model. When `RUN-H8T6` shows up, the script ran. Each can happen without the other.

What varied was how `SKILL.md` pointed at the files:

`For the complete formatting rules, see [reference.md](reference.md)`, `For a finished entry, see [examples/sample-entry.md](examples/sample-entry.md)` and `To check an entry before you reply, run [scripts/check-entry.sh](scripts/check-entry.sh) with the bullet as its only argument`.` SKILL.md`.`@reference.md` and `@${CLAUDE_SKILL_DIR}/examples/sample-entry.md` in place of the first two links, with the script line from link.
Both prompts named the skill, so whether the model would pick it on its own was not part of the test ([a separate article](https://dev.to/rulestack/will-claude-code-call-your-skill-on-its-own-38-runs-one-skill-seven-ways-to-describe-it-203m) measured that). The first asked for an entry for "the parser no longer crashes when the input file is empty". The second described a breaking change to `--config`, which should set off the `BREAKING:` rule in `reference.md`.

The 20 runs: link, must and none with each prompt, twice each (12 runs); must with the first prompt and `reference.md` padded to 25,405 characters by a 300-line appendix (2); at (2); at with an unrelated `reference.md` placed at the project root as a decoy (2); and link with the first prompt on Sonnet (2). Every run used this command line from the project root, with `--model sonnet` for the Sonnet pair:

```
CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 claude -p "$PROMPT" \
  --output-format stream-json --verbose --model opus \
  --permission-mode default --allowedTools "Bash" --max-turns 10 \
  --settings '{"disableAllHooks": true}' \
  --setting-sources project --strict-mcp-config
```

`--allowedTools "Bash"` pre-approved shell commands so that no run stopped at a permission prompt. All 20 runs ended with `subtype: success`, no permission denials and no failed tool calls. Together they cost $1.12.

If a supporting file cost something "every time", it would show up in the first request of the session, so I compared that request across variants. All 12 Opus runs with the first prompt sent a first request of exactly 19,152 tokens, counting input, cache reads and cache writes. Those 12 runs covered four `SKILL.md` files between 543 and 936 bytes, `reference.md` at 1,182 and at 25,405 characters, and the decoy file at the project root. The six Opus runs with the second prompt all sent 19,173 tokens, and both Sonnet runs 19,158. No body and no supporting file moved the count by a single token.

The transcripts agree. Apart from my prompt naming it, the only trace of the skill before the first response is a `skill_listing` attachment whose content is one line: "- changelog-entry: Writes one changelog entry for a code change in this repository. Use when the user asks for a changelog entry, a release note line, or a CHANGELOG update." None of the markers appears before the first response in any of the 20 transcripts. The docs say as much: "Unlike CLAUDE.md content, a skill's body loads only when it's used, so long reference material costs almost nothing until you need it."

In all 20 runs, the model's first action was a `Skill` call. Its tool result is one line, "Launching skill: changelog-entry", and Claude Code then adds a single user message, flagged as meta, that holds the rendered body. The frontmatter is gone, and the first line is not in `SKILL.md` at all:

```
Base directory for this skill: /…/proj/.claude/skills/changelog-entry

# Changelog entry

Write one entry for the change the user describes, for the Unreleased section of the project's changelog.
…
```

When the model passed the change description as the skill's arguments, which it did in 13 of 20 runs, an `ARGUMENTS: …` line followed the body. That is the fallback for bodies without placeholders, covered in [the arguments article](https://dev.to/rulestack/claude-code-skill-arguments-across-12-runs-what-arguments-0-and-argument-hint-actually-put-in-1881).

The skills page does not mention the "Base directory" line, but it is what makes the docs' relative links work. The `Read` tool's description says "`file_path` must be an absolute path", and all 33 `Read` calls in the 20 runs used an absolute path under that directory. The Bash commands reached the same directory through `cd` or an absolute path, except one that used `cd .claude/skills/changelog-entry` from the project root. None of them failed.

The invocation step is the second place where a supporting file could have slipped in early. In the 16 runs without `@` lines, the request after the `Skill` call grew by 328 to 558 tokens, depending on the length of the body and on whether arguments were passed. The must runs with the first prompt grew by 514 tokens with arguments and 452 without, with the normal and with the padded `reference.md` alike. Invoking the skill brought in the body and nothing from the supporting files, as documented.

Here are all 20 runs by pointer style. "In context" means the file's marker appears in a tool result or an attachment in the transcript.

| Pointer in `SKILL.md` | Model | Runs | `reference.md` in context | example in context | script executed | script source in context | 
|---|---|---|---|---|---|---|
| markdown links (docs pattern) | Opus | 4 | 4 | 4 | 4 | 4 | 
| markdown links (docs pattern) | Sonnet | 2 | 2 | 2 | 2 | 1 | 
| "you must read / run" | Opus | 6 | 6 | 6 | 6 | 4 | 
| `@` references | Opus | 4 | 4 | 4 | 4 | 4 | 
| none | Opus | 4 | 0 | 0 | 0 | 0 | 

The must row includes the two padded runs. With any pointer, `reference.md` and the example reached the model in 16 of 16 runs, no later than the first step after invocation: as parallel `Read` calls, as one `cat` of several files, or, for the example in three at runs, as an attachment that arrived with the body. Without a pointer, nothing was read in 4 of 4 runs. The none runs did not list the skill directory, did not search and did not open a file. Each finished after two requests, in 3.1 to 3.7 seconds, with an answer built from the body alone.

The answers show the price of that. The train tag from `reference.md` appears in 16 of 16 answers from runs with a pointer and in 0 of 4 from the none runs, whose entries the release script described in `reference.md` would reject. For the breaking `--config` change, 4 of 4 pointer answers had `BREAKING:` after the area tag and 0 of 2 none answers did. All four none answers also put a period before the entry code, which `reference.md` rules out. A supporting file that `SKILL.md` did not mention was not "loaded when needed". It was not loaded at all, because nothing told the model that it existed.

The other side is that no pointer was treated as optional. Even for a one-line bug fix, every run with a pointer read the full rules and the example, so with these pointers "only when needed" meant "every time the skill runs". Whether a pointer that states a condition, such as "read reference.md only for breaking changes", makes the read conditional, I did not test.

The answers also show why I did not rely on them. In six runs the entry left out the example's footer line, and five of those answers explained the choice, for example "I left that out because it looks tied to the sample's `[cli]` change, not a rule for every entry." In two runs whose transcripts show the example being read, the example's marker appears nowhere in the answer. The script's source marker, `SRC-W5N3`, is in none of the 16 answers. No answer could have told me whether the script had been read.

All 16 runs with a pointer executed `scripts/check-entry.sh`. `RUN-H8T6` appears in a Bash tool result in every one of those transcripts, and 16 of 16 answers carry the `Checked: RUN-H8T6` line the script asked for. In 13 of those 16 runs, `SRC-W5N3` appears in the transcript too, which means the script's source went into the model's context. The 13 runs got there in three ways.

`Read` of `scripts/check-entry.sh` in the same step as their other reads, then ran the script in the next step.`cat` in an earlier Bash call, one to print all three files at once, the other right after `ls -R` of the skill directory.
The model's own descriptions of those combined commands show that printing the source was intended: "Show check script and validate entry", "Show and run the entry checker script" and "Show and run the entry check script". The fourth, from a Sonnet run, had no description.

Only 3 runs executed the script without reading it: one must run with the normal `reference.md`, one with the padded file, and one Sonnet run. By pointer style, the source was loaded in 4 of 4 Opus link runs, 4 of 6 must runs, 4 of 4 at runs and 1 of 2 Sonnet link runs. Two runs per configuration are too few to rank the styles against each other. What the numbers do show is that neither the docs' link pattern nor a plain "you must run" kept the source out.

Both script pointers used the wording the best-practices page gives for execution, "run" plus the script's name, and neither said "see" or "read". Neither gave a complete command line, so the model had to work out the invocation, and the script's fourth line is a usage comment. The transcripts store the model's reasoning as empty text, so I cannot tell whether it opened the script to find the usage or to look at code it was about to run. I did not test a pointer that gives the exact command, like the `python3 ${CLAUDE_SKILL_DIR}/scripts/visualize.py .` line in the docs' own codebase-visualizer example.

For a 19-line script the price was small, as the next section shows. The point is a different one: a script is not a place to keep anything away from the model, and the saving the docs give as the reason to bundle code as scripts, not having the code in context, held in 3 of these 16 runs.

Each request's size in a transcript is `input_tokens + cache_read_input_tokens + cache_creation_input_tokens`, and the growth from one request to the next is what the step in between added. Runs that made the same tool calls grew by the same amount, to the token:

| The step read | Runs | Next request grew by | Model output in that step | File contents (the difference) | 
|---|---|---|---|---|
| `reference.md` + example | 3 | 1,020 | 313 | 707 | 
| `reference.md` + example + script | 4 | 1,568 | 472 | 1,096 | 
| padded `reference.md` + example | 2 | 11,622 | 313 | 11,309 | 

The model's output in that step is its two or three `Read` calls, which are long because the lab's absolute path is long. The difference is what the `Read` tool returned: the file contents, with line numbers. The script's source added 389 tokens to the results, and its extra `Read` call 159 more. Padding `reference.md` by 24,223 characters added 10,602 tokens, and every run that pointed to the file paid them.

Across whole runs, the four none runs ended with a final request of 19,480 to 19,607 tokens and cost $0.011 to $0.014 each. The 12 Opus runs with a pointer and normal-size files ended at 21,467 to 22,092 tokens, took 14.8 to 31.7 seconds and cost $0.042 to $0.084. They made four requests instead of two and wrote 714 to 1,600 output tokens instead of 88 to 169. The two padded runs ended at 31,849 and 32,134 tokens and cost $0.134 and $0.141, against $0.055 and $0.057 for the same must runs with the normal file.

`@` lines load at invocation, from the working directory
The skills page mentions `@` file references only once, in the section on skills synced from a claude.ai account, where it says that in some sessions Claude Code "doesn't attach the files that `@` references name the way it does for a local skill". The at variant shows what that means for a local skill.

`@${CLAUDE_SKILL_DIR}/examples/sample-entry.md` worked as a preload in 3 of 4 runs. Claude Code substituted the variable and attached the example in the same step as the body, so the model had it before making a single call. The transcript records it as an attachment of type `file` holding the path and the full content. It added 278 tokens to the invocation step: 804 in the two at runs without arguments, against 526 in the one run without arguments where nothing was attached.

`@reference.md` never attached the skill's `reference.md` (0 of 4). It resolved against the working directory, here the project root, not against the skill's directory the way the model had resolved the markdown links. In the two at runs without a `reference.md` at the root, nothing was attached for that line. In one decoy run Claude Code attached the unrelated root file, API notes for a TOML library carrying the marker `DCY-R2L7`. The model then listed the skill directory, printed the skill's own `reference.md` and script, and followed the right rules. In the other decoy run nothing was attached at all, neither the decoy nor the example, and the model read all three skill files with `Read`. I could not find out why the two decoy runs differed.

In all 4 at runs the model still read the skill's own `reference.md`, because "see @reference.md" also worked as a plain pointer. So an `@` line with `${CLAUDE_SKILL_DIR}` can put a file into context together with the body, and a bare relative `@` line can put the wrong file there.

Point to every supporting file from `SKILL.md`. In these runs a file that was not mentioned was never found, and the model listed the skill directory only once, after an `@` line had attached the wrong file.

Budget a pointed-to file as if it loads on every invocation. The docs' link pattern was followed in 6 of 6 runs and the "must" wording in 6 of 6, even for a one-line fix. If a file should load only for some requests, write the condition into the pointer and test it, because I did not.

Do not treat `scripts/` as a place the model will not look. If a script's source must stay out of context, because it is large or because the model should not see it, I would try an exact command line with `${CLAUDE_SKILL_DIR}`, as in the docs' visualizer example. For a script that needs no input from the model, the `!` injection syntax is another candidate: the docs say it "runs shell commands before the skill content is sent to Claude" and that "Claude receives actual data, not the command itself". Both are untested here.

Use `@${CLAUDE_SKILL_DIR}/...` when a file must arrive together with the body, and avoid a bare `@file` in a skill.

Check the transcript, not the answer. The script's source marker was in 13 of 16 transcripts and in 0 of 16 answers.

Twenty runs are a small sample: two per configuration, one skill, one tiny project, two models. The per-style split of script reads could move a lot with more runs. The 13 of 16 total is the number I would stand behind, not the ranking of pointer styles.

Every run was headless, with Bash pre-approved and a prompt that named the skill. I did not test an interactive session, where running the script may stop for a permission prompt, a skill typed as a slash command, `context: fork`, plugin or synced skills, or subagents with preloaded skills, for which the docs say "the full skill content is injected at startup".

I did not test pointers with an exact command line, `${CLAUDE_SKILL_DIR}` in the script path, `!` injection, conditional pointers, or references from `reference.md` to further files, which the best-practices page warns Claude may preview with `head -100` instead of reading in full. The script had 19 lines, and I did not test whether a long script gets read the same way.

Each session held one task. I did not test a second request in the same session, compaction, or whether a file read once gets read again.

The model's reasoning is stored as empty text in the transcripts. The result events report 139 to 702 thinking tokens for each run with a pointer, but not their content, so why the model read the script is inferred from its tool calls and their descriptions, not observed.

The difference between the two decoy runs is unexplained. The only lead I found is a one-second abort timer that strings in the 2.1.285 binary show around attachment collection, and I did not check whether it fired.

Everything ran on macOS.

A cut-down version of the lab fits in one paste. `SKILL.md` and the script below are byte for byte the files I tested with the link pointer. `reference.md` and the example are shortened, so treat your result as a new sample rather than a replay.

```
D=$(mktemp -d) && cd "$D"
S=.claude/skills/changelog-entry
mkdir -p "$S/examples" "$S/scripts"
cat > "$S/SKILL.md" <<'EOF'
---
name: changelog-entry
description: Writes one changelog entry for a code change in this repository. Use when the user asks for a changelog entry, a release note line, or a CHANGELOG update.
---

# Changelog entry

Write one entry for the change the user describes, for the Unreleased section of the project's changelog.

- One bullet, in the past tense, starting with the area in square brackets, such as `[parser]`.
- End the bullet with the entry code `(BODY-M3V8)`.
- Reply with the finished entry only. Do not create or edit any file.

## Additional resources

- For the complete formatting rules, see [reference.md](reference.md)
- For a finished entry, see [examples/sample-entry.md](examples/sample-entry.md)
- To check an entry before you reply, run [scripts/check-entry.sh](scripts/check-entry.sh) with the bullet as its only argument
EOF
cat > "$S/reference.md" <<'EOF'
# Changelog entry rules (complete)

Every bullet ends with two codes, in this order: the entry code from SKILL.md, then the release-train tag `(train REF-Q7X4)`.
The release script rejects any bullet that does not carry the train tag.
EOF
cat > "$S/examples/sample-entry.md" <<'EOF'
# Sample entry

Copy this layout, including the indented footer line under the bullet:

- [cli] Fixed `--verbose` being ignored when combined with `--quiet` (BODY-M3V8)
  Changelog-Set: EXM-K2P9
EOF
cat > "$S/scripts/check-entry.sh" <<'EOF'
#!/bin/sh
# check-entry.sh: checks one changelog bullet against the house rules.
# Source marker: SRC-W5N3
# Usage: check-entry.sh "<the bullet line>"
entry="$1"
if [ -z "$entry" ]; then
  echo "usage: check-entry.sh \"<bullet>\"" >&2
  exit 2
fi
case "$entry" in
  "- ["*) ;;
  *) echo "FAIL: the bullet must start with '- [area]'"; exit 1 ;;
esac
words=$(printf '%s' "$entry" | wc -w | tr -d ' ')
if [ "$words" -gt 30 ]; then
  echo "FAIL: $words words; keep the bullet shorter"; exit 1
fi
stamp=$(printf 'RUN-%s%s' 'H8' 'T6')
echo "OK. Add this footer line under the bullet: Checked: $stamp"
EOF
chmod +x "$S/scripts/check-entry.sh"
claude -p 'Use the changelog-entry skill to write the changelog entry for this change: the parser no longer crashes when the input file is empty.' \
  --output-format json --model opus --permission-mode default \
  --allowedTools "Bash" --max-turns 10 --settings '{"disableAllHooks": true}' \
  --setting-sources project --strict-mcp-config > out.json
```

Then find the transcript, see where each marker first appears, and list the size of every request:

```
SID=$(jq -r .session_id out.json)
T=$(ls ~/.claude/projects/*/"$SID".jsonl)
for m in BODY-M3V8 REF-Q7X4 EXM-K2P9 SRC-W5N3 RUN-H8T6; do
  printf '%s ' "$m"
  jq -r --arg m "$m" 'select(.type=="user" or .type=="attachment") | select(tostring | contains($m)) | if .type=="attachment" then "attachment:" + .attachment.type elif (.message.content|type)=="array" then "user:" + .message.content[0].type else "user:text" end' "$T" | head -1
  echo
done
jq -r 'select(.type=="assistant") | [.requestId, (.message.usage | .input_tokens + .cache_read_input_tokens + .cache_creation_input_tokens)] | @tsv' "$T" | uniq
```

`SRC-W5N3` followed by `user:tool_result` means the script's source reached the model. `RUN-H8T6` without `SRC-W5N3` means the script ran unread, which is what happened in 3 of my 16 runs.

*Rulestack writes skills, slash commands and rules files for Claude Code and sells them at [rulestack.gumroad.com](https://rulestack.gumroad.com?utm_source=devto&utm_medium=article&utm_campaign=13-of-16-claude-code-runs-loaded-the-skill-script-the-docs-call-executed-not-loaded). The marker trick from this article, one unique string per file plus a stamp that only the script's output contains, takes a few minutes to add to any skill you want to audit before you ship it.*

*If one of your runs executes a bundled script without its source showing up in the transcript, especially with an exact command line in `SKILL.md`, leave the Claude Code version and the pointer wording in the comments below, and follow [@ai-shop.bsky.social](https://bsky.app/profile/ai-shop.bsky.social) for the next measurement.*
