The company slashed its coding agent's system prompt from roughly 800 tokens to 164, reflecting a philosophical shift in how advanced AI models should be instructed
Anthropic just did the AI equivalent of Marie Kondo-ing its codebase. The company reduced the system prompt powering Claude Code by approximately 80%, dropping from around 800 tokens down to just 164. The kicker: performance didn’t budge. If anything, it may have gotten better.
The change was announced on July 2, 2026, by Anthropic engineer Thariq Shihipar at the AI Engineer World’s Fair.
Less is literally more #
For the uninitiated, a system prompt is the set of instructions that tells an AI model how to behave before a user ever types a word. Claude Code’s old system prompt was roughly 800 tokens of detailed instruction. The new prompt is 164 tokens. Anthropic gutted more than three-quarters of the instructions and the model kept performing at the same level on benchmarks, or potentially improved. Claude Code is a command-line tool that lets developers manage coding tasks through conversation with an AI agent. Every time a developer sends a request, the system prompt tags along for the ride. Cutting that prompt by 80% means fewer tokens consumed per interaction, which translates directly to lower costs and faster response times for every single user.
The philosophy behind the trim #
The reduction wasn’t just a cleanup exercise. It reflects a deeper shift in how Anthropic thinks about instructing its models.
According to Shihipar, the company’s newer Fable 5 (Mythos-class) models demonstrate something counterintuitive: they actually work better with less direction. These models show enhanced imaginative capabilities, and overly prescriptive prompts can actively stifle their performance.
Earlier generations of AI models relied on explicit instructions, worked examples, and carefully defined boundaries to produce useful results. The newer architecture has internalized enough understanding that it can handle ambiguity and context without having every rule spelled out in advance.
This represents a broader strategic pivot at Anthropic away from rigid, hard-coded prompts toward a context-driven approach. Rather than maintaining one massive, universal system prompt, the new architecture employs specific, targeted prompts tailored to individual models.
What this signals for the industry #
Lower token consumption per request means Anthropic can offer Claude Code at more attractive price points, or maintain current pricing while improving margins. For developers running Claude Code at scale across large codebases, even small per-request savings compound quickly into meaningful cost differences.
Shorter system prompts leave more room in the context window for actual user content. Freeing up roughly 636 tokens per request gives the model more breathing room to focus on the user’s problem.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our