Claude has a persistent habit of wrapping every technical answer in layers of "slop"—metaphors, repetitive recaps, and those irritating "it's not X, it's Y" linguistic patterns. Even when using the native concise mode, the output often retains too much filler. System prompts are a temporary fix at best, as they typically degrade or are ignored entirely after a few turns in a long conversation. The Terse plugin solves this by forcing the model to strip away the fluff and get to the point immediately.
How does Terse actually reduce output length? #
The plugin targets the specific linguistic quirks that make LLM responses feel bloated. Instead of just asking for "brevity," which the model often interprets as "summarize this," Terse focuses on the removal of conversational filler. This means eliminating the introductory slogans and the concluding summaries that add no value to a developer's workflow.
When a model ignores a system prompt, it is usually because the conversational context has grown too large, and the "concise" instruction loses weight compared to the established pattern of the chat. By implementing this as a plugin for Claude Code, the constraint is applied more consistently across the session.
Implementing the Terse constraint #
To get Claude to stop the "essay" behavior, you need a prompt that explicitly bans specific structural patterns rather than just requesting a shorter length. The following prompt structure is what drives the logic behind the Terse approach:
Act as a terse technical assistant.
1. Eliminate all introductory filler (e.g., "Sure, I can help with that," "Here is the solution").
2. Remove all concluding summaries or "hope this helps" phrases.
3. Avoid metaphors, slogans, and contrast-style explanations (e.g., "It's not just X, it's Y").
4. Provide the direct answer or code block immediately.
5. If a one-word answer suffices, use only one word.
When should you use a plugin over a system prompt? #
If you find that Claude reverts to long-winded explanations after five or six exchanges, a standard system prompt is failing you. This happens because the model prioritizes the "persona" it has built during the current session over the initial instructions.
A dedicated plugin approach is necessary when:
- You are doing heavy coding where every extra line of text increases the time spent scrolling.
- You are using the CLI version of the tool where screen real estate is limited.
- You find yourself manually deleting "Here is the updated code" from every single response.
Dealing with the "Concise Mode" failure #
Many users rely on the built-in concise toggle, but that often fails because it reduces the number of points made without actually changing the style of the prose. You still get the same polite, corporate tone, just in a smaller paragraph. Terse is different because it attacks the prose style itself.
The failure point for most "shorten this" prompts is that they don't define what "filler" is. By explicitly banning "recaps" and "strange choice of nouns," the output becomes purely functional. If the output is still too long, the next step is to move from "concise" instructions to "negative constraints"—telling the model exactly what not to write.
The goal is to move the interaction from a conversation to a data exchange. When the model stops trying to be a "helpful assistant" and starts acting like a precise tool, the development velocity increases because you spend less time filtering noise and more time implementing code.
Next Struggling with location-based blocks on Google Antigravity IDE →
All Replies (2) #
Want a live back-and-forth? Join the global AI chat room — login to talk.