cd /news/ai-agents/whoa-my-claude-code-skill-descriptio… · home › topics › ai-agents › article
[ARTICLE · art-148246] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Whoa… My Claude Code Skill Descriptions Are Way Too Long…?!

A developer audited 54 personal Claude Code skills and found their descriptions totaled 28,131 characters, with a median of 435.5 characters and 10 exceeding 800 characters. Running Claude Code's /skill-doctor diagnostic revealed that skills with long descriptions and zero usage history were listed by name only in the system prompt, because the skill listing budget is allocated by a usage score that multiplies use count by a factor halving every seven days since last use. The developer traced the behavior to Claude Code 2.1.295's source, where over-budget listings drop whole descriptions rather than truncating them, and rewrote the descriptions accordingly.

by read9 min views3 publishedOct 9, 2026

I was looking over the descriptions of my skills and thought, "Aren't these too long?" Before the rewrite, I had 54 skills in ~/.claude/skills/. Their descriptions totaled 28,131 characters, the median was 435.5 characters, and 10 of them were over 800 characters.

To see what long descriptions do, I ran Claude Code's /skill-doctor (a diagnostic command that shows, for each loaded skill, how much context it uses every turn). Here is part of the output.

  skill                                              source                    context  7d tokens   uses  last used
  jev-judgment-design                                userSettings                 < 20          -     0×  never
  loop-design-check                                  userSettings                 < 20          -     0×  never
  mono-color                                         userSettings                 < 20          -     0×  never
  repair-discipline                                  userSettings                 < 20          -     0×  never
  …
  context = this skill's one-line listing in the system prompt, included every turn

context is the size of the one line the skill gets in the skill listing sent every turn. These four skills have descriptions of 459 to 1,048 characters. Yet each is listed at under 20 tokens, so the listing carries only their names. All four were skills I had never used. Among the skills I wasn't using, the longer I had written a description, the more likely the model saw none of it.

This article checks why that happens against Claude Code's implementation and shows how I rewrote my descriptions. Whether that is the right way to write them, I don't know yet.

Every turn, Claude Code puts a listing of the skills the model can invoke on its own into the system prompt. Each line has the form - name: description. The listing includes not only the skills I wrote but also plugin skills and skills synced from claude.ai. In my environment, 94 skills are loaded, and all of them compete for one listing budget.

The budget counts display width, not characters. Width is measured with Bun.stringWidth, so a half-width character counts as 1 and a full-width character, such as a Japanese one, counts as 2. In the code of Claude Code 2.1.295 itself, the budget was this formula:

skillListingBudgetFraction) When the listing goes over budget, Claude Code writes a "Skill listing over budget" warning to the debug log and chooses which skills get a description:

When the budget runs short, no description is cut off midway (only the part over the per-skill cap of 1,536 characters is cut). A skill either gets its whole description or gets only its name.

The usage score formula looks like this in the code (verbatim, still minified):

function xwt(e){let n=Oa(ce().skillUsage??{},e);if(!n)return 0;let r=(Date.now()-n.lastUsedAt)/86400000,s=Math.pow(0.5,r/7);return n.usageCount*Math.max(s,0.1)}

It is the use count multiplied by a factor that halves every 7 days since the last use, with a floor of 0.1. A skill you have never used scores 0.

The four skills from the opening had a use count of 0, so their score was 0 and they came last. Their long descriptions didn't fit in the remaining budget, so they were listed by name only. From the model's side, a name-only skill gives no clue what it does. If it isn't invoked, its score stays 0, and as long as the budget stays short, it is listed by name only in the next session too. This is a loop I read from the code; I have not counted invocations to check whether they were invoked less often.

Even among skills with the same usage score of 0, whether one ends up as a name only depends on the budget left at that point and the length of its description.

This wasn't the first time I'd noticed my descriptions were long. At the end of August, I had also tried to cut down how much of them sat in the system prompt every turn. As of August 30, the skills in my repository (excluding two placed by symlink from outside repositories) numbered 57, and their descriptions totaled 26,738 characters. Just before this rewrite, counted the same way, there were 52 skills and 27,105 characters. I had five fewer skills and more characters.

Every time I add a skill or fix one, its description gains another "also use it when…". Even when I cut the number of skills, each remaining skill's line keeps getting longer.

I took the direction of the rewrite from the design guidelines of Matt Pocock's skill collection. He splits skills into two by who invokes them (invocation) (invocation.md).

/name gets disable-model-invocation: true, and its description is a one-line summary for humans On how to write descriptions for the model, another of his skills (writing-for-agents) calls the description a "context pointer" to the body, and says to write one trigger per branch and to phrase things positively rather than negatively. Writing them short was my own call, made after looking at the budget.

Claude Code's implementation matches this split: skills with disable-model-invocation: true are removed before the skill listing is built. Since they aren't in the listing, they use none of the budget. I moved two skills to this side. One was grill-me, which interrogates a plan one question at a time: it only ever ran when I asked for it.

Before: A relentless one-question-at-a-time interview that stress-tests a plan or design before you build. Use when the user wants to pressure-test a plan, says "grill me" / "grill this" / "stress-test this" / "poke holes in this" / "interview me about this design", or before committing to expensive or hard-to-reverse implementation work while the goal is still vague.
After:  A relentless one-question-at-a-time interview that stress-tests a plan or design.   (+ disable-model-invocation: true)

Skills the model invokes appear in the listing every turn. So I write their descriptions as pointers that stay resident every turn. The description says only when to use the skill; the steps and the output shape go in the body. The one that shrank the most is measurement-discipline, a skill for making decisions from measured results.

The Before listed the same branches as 10 pairs of Japanese and English phrasings, and named three situations not to use it in (1,046 characters).

Discipline for designing or evaluating measurement-based claims, thresholds, guards, experiment results and observation periods. Use when the user says 「この実験結果で判断していい?」 / "can I decide on this experiment result?", 「閾値を決めたい」 / "I want to set a threshold", 「ガード/検査を足したい」 / "I want to add a guard or check", 「1 回通ったから大丈夫」 / "it passed once, so it's fine", 「観察期間はどれくらい」 / "how long should the observation period be?", 「いつゲートを開く」 / "when do we open the gate?", 「shadow のまま何週待つ」 / "how many weeks do we wait in shadow?", 「この RFC 塩漬けでは」 / "isn't this RFC just sitting idle?", 「本番より良い」 / "it's better than production", 「本番と比べて」 / "compare it against production", when a design places a numeric threshold, a suspicion flag, or a wait-for-N-observations condition, when a candidate is compared against production, or when a claim rests on measured data. NOT for — designing the instruments themselves (read-only distributions and readings) — out of scope here; designing an LLM judge (llm-as-judge); whether a loop's structure is sound (out of scope here).

The After has one trigger per branch, and puts the words I only ever type in Japanese in parentheses, once each (291 characters).

Measurement discipline for claims that rest on data. Use when deciding on an experiment result, setting a threshold (閾値) or guard, choosing an observation period (観察期間) or when to open a gate, comparing a candidate with production (本番比較), or when an automated gate rejects outputs you trust.

A Japanese character takes a width of 2, so writing both Japanese and English pays the budget twice. In this example, I removed every situation not to use it in. I keep a "not for" only where a neighboring skill competes for the same trigger, and I write it in the positive form: "for X, use y".

Only the skills on the right use the listing budget. The reason to write a description short for the budget exists only on the right.

Since trimmed descriptions grow back, being careful while writing isn't enough. I added a lint rule that counts listing width to my harness's pre-commit checks.

LISTING_WIDTH_MAX = 12_000
DESCRIPTION_WIDTH_MAX = 400

def listing_width(text: str) -> int:
    """Claude Code が listing 行を測る Bun.stringWidth の近似 — 全角 (W/F) を 2 と数える。"""
    return sum(2 if unicodedata.east_asian_width(ch) in "WF" else 1 for ch in text)

It measures the listing line of each model-invoked skill, counting a full-width character as 2, and blocks the commit when one skill goes over 400 or the total goes over 12,000. 12,000 is 40% of the 30,000 budget for a 1M window. I chose 400 and 12,000 myself; I did not calibrate them against anything.

Here is the result of the rewrite.

Before After
Total description length (54 skills, characters) 28,131 10,378
Median (characters) 435.5 189.5
Skills over 800 characters 10 0
Listing width of model-invoked skills 24,203 9,780
Skills in ~/.claude/skills/ under 20 tokens in/skill-doctor 4 0

In /skill-doctor after the rewrite, the context of the four skills jev-judgment-design, loop-design-check, mono-color and repair-discipline went from < 20 to between ~80 and ~250: they got their descriptions and appear in the listing with them. All four still have a use count of 0.

I have not measured whether the rewrite changed how often the model invokes skills on its own.

I do have an experience in the opposite direction. Half a year ago, I revised the wording of the description of a skill called search-first. The share of times the model invoked it on its own dropped from 27% to 8%, and I reverted the change (the denominator isn't in my records). All I know is that I can't say a polished description gets a skill invoked more.

I decided three things: to settle invocation first, to write the model-invoked skills short, and to count width with a lint rule. With that, none of my own skills is listed by name only anymore. But a description being visible and a skill being invoked when it's needed are two different things. What should a skill description be? I don't have that answer yet.

disable-model-invocation and other fields How this was written: Claude (Claude Code) wrote the prose and translated it into English. It worked from my session records, measurements of the descriptions before and after the rewrite, and conversations with me. I decided to settle invocation first and then rewrite, and to hand over the question rather than an answer. I am responsible for its content.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/whoa-my-claude-code-…] indexed:0 read:9min 2026-10-09 · —