{"slug": "give-your-llm-more-dsls", "title": "Give your LLM more DSLs", "summary": "A blog post argues that LLMs should be given domain-specific languages (DSLs) to edit rather than free-form prompts, citing a captain-hook example where Claude Code hooks in nested JSON are replaced by a few lines of DSL such as block_command([\"rm\"], reason=\"rm is banned in this repo\", hint=\"use trash instead\"). The post also describes Aneta's context compiler, which represents data, prompts and examples in a unified IR, runs cache-aware optimization passes, and renders provider-specific output; for one request, no passes sent 9,464 tokens with 8,564 read from cache, COST sent 8,804 with 8,564 cached, QUALITY sent 6,566 with 4,980 cached, and SPEED sent 2,386 with 680 cached.", "body_md": "# Give your LLM more DSLs!\n\n+ Ask once, compile the answer, run many times\n\nIt’s been thoroughly shown at this point that having LLMs author prompts for themselves (free-form that is, not DSPy-style) is disastrous, and that [LLM-authored skills actually make things worse](https://arxiv.org/html/2602.12670v4), but [writing code is what these next-token-predictors excel at](https://arxiv.org/abs/2402.01030)!\n\nThere’s a frequently-used trick for getting LLMs to do out-of-band tasks, which is: come up with a thoughtful DSL, capture your current state as a DSL, get the LLM to manipulate the DSL, then translate that back to your domain. It turns out this works pretty well for *in*-domain tasks too. A great example of this is the continuous improvement loop in [`captain-hook`](https://yasyf.com/writing/less-prompts-more-guardrails/): getting the LLM to scan transcripts for corrections the user made, then tracking and codifying those using a DSL over Claude Code hooks gives you a nice anti-slop flywheel.\n\nyour hooks live in nested json, but the model only ever edits a few lines of dsl.\n\n```\n{\n  \"hooks\": {\n    \"PreToolUse\": [\n      {\n        \"matcher\": \"Bash\",\n        \"hooks\": [\n          {\n            \"type\": \"command\",\n            \"if\": \"Bash(rm *)\",\n            \"command\": \"echo 'BLOCKED: rm is banned in this repo. use trash instead.' >&2; exit 2\"\n          },\n          {\n            \"type\": \"command\",\n            \"if\": \"Bash(pnpm install*)\",\n            \"command\": \"echo 'BLOCKED: pnpm install at the root. cd into the workspace first.' >&2; exit 2\"\n          }\n        ]\n      }\n    ]\n  }\n}\nblock_command(\n    [\"rm\"],\n    reason=\"rm is banned in this repo\",\n    hint=\"use trash instead\",\n)\nblock_command(\n    [\"pnpm\", \"install\"],\n    reason=\"pnpm install at the root\",\n    hint=\"cd into the workspace first\",\n)\n```\n\n- added\n\nAnother great example came up recently while building Aneta’s context compiler.\n\nThe idea behind the context compiler is pretty simple:\n\n1. Take a bunch of data, prompts, and examples and represent them in a unified IR.\n2. Run cache-aware optimization passes over that IR (every pass weighs the tradeoffs of busting the cache vs decreasing tokens, depending on what that particular call is optimizing for).\n3. Render that IR to the provider-specific format for the LLM being called.\n\nThe trick is capturing a closure at IR-construction time that knows how to render a given piece of context in both verbose and compact ways, and then allowing the optimization passes to toggle one or the other (in addition to many standard compaction techniques). The compiler can be optimizing for cost, quality, speed, or some linear combination of those.\n\nthe same context compiled for `COST`, `QUALITY`, and `SPEED`, where each pass weighs the tokens it saves against the cached tokens it gives up.\n\ncontext nodes, in the order they are senttokens\n\nopen a node to see both of its renderings.\n\n- cache ends here\n- verbose, 900 tokens \n\n``` php\nquery_table(arm=\"b\") ->\n{\"arm\": \"b\", \"n\": 12, \"index\": 0.44,\n \"sd\": 0.03, \"p\": 0.01,\n \"rows\": [12 rows of raw counts]}\n```\n\n compact, 240 tokenssent \n\n```\n{\"arm\": \"b\", \"n\": 12, \"index\": 0.44,\n \"sd\": 0.03, \"p\": 0.01}\n```\n\n- read from cache\n- sent fresh\n- cut by a pass\n\npasses for `COST`\n\n- runscompact the new tool resultsaves 660 tokens, loses none from cache\n- skipscompact the examples and the tablesaves 3,520 tokens, loses 7,884 from cache\n- skipsswap figures 1 and 2 for their 24-token captions, since this turn needs only figure 3saves 2,898 tokens, loses 3,584 from cache\n\n`COST` only runs a pass that leaves the cache alone, since a cached token bills at a fraction of a fresh one.\n\n| tokens in this request |  |  | \n|---|---|---|\n| objective | total | read from cache | \n|---|---|---|\n| no passes | 9,464 | 8,564 | \n| `COST` | 8,804 | 8,564 | \n| `QUALITY` | 6,566 | 4,980 | \n| `SPEED` | 2,386 | 680 | \n\nCOST sends 8,804 tokens, 8,564 of them read from cache.\n\nOne of the more interesting optimization passes we do is around images. Images take up HUGE numbers of tokens, yet are often necessary to include in scientific queries as they are figures that are referenced, and OCRing them just doesn’t give sufficient fidelity. However blindly including them results in much larger contexts, and the quality degradation that comes with those. When operating in `SPEED` or `QUALITY` modes, the compiler has a pass that allows it to selectively exclude images from the context when compiling (along with a tool for the LLM to call if it then decides it needs that image). Except, how do you decide when to include the image?\n\nwhat re-sending every figure on every turn costs, against sending only the one each turn needs.\n\none figure, cut into 28 px patches\n\n1000 x 1000 px -> 36 x 36 patches -> 1,296 tokens\n\nover a thread\n\nimage tokens sent over 9 turns\n\n- re-send every figure\n- 23,328 tokens\n- send only the needed one\n- 11,664 tokens\n\nre-sending every figure costs 2x.\n\nYou can start with simple heuristics like how long ago the image was last mentioned, but that breaks down pretty fast in deep research. Another idea is to have a small sidecar LLM (Haiku-class or smaller) make the decision at every turn; this works surprisingly well when in `QUALITY` mode, but is prohibitively slow for `SPEED`. Inspecting the decisions that sidecar is making yields an interesting insight however: given a static summary of the conversation the decision is a pure function of the message about to be sent.\n\nshould the chart be included with each message? the right call first, then what each strategy did.\n\n- right call\n- in when answering the message needs the chart\n- heuristic\n- in when the message or one of the 2 before it says \"chart\"\n- sidecar\n- a small llm decides on every turn\n\n- wrong call\n- still deciding\n\neach sidecar call holds its message for about 4,100 ms, shown here 5x faster.the heuristic answers in under 1 ms and gets 3 of 10 wrong. the sidecar gets all 10 right, but each message waits about 4,100 ms for it.\n\nSo the question we want to quickly answer is: given a somewhat-recent summary of the conversation, and the message about to be sent, should we include this image? Most of the features the sidecar LLM gathers to make that decision end up being either those easily derived from classic NLP toolboxes, or small binary decisions that an on-device model could make. So how do we handle image inclusion in `SPEED` mode? You can probably guess where I’m going with this… a DSL!\n\nIt turns out you can just ask the same sidecar model to write a small function, and provide it with a few helpers. Here are the ones we provide:\n\n``` php\nasync def nli_scores(a: str, b: str) -> tuple[float, float, float]: ...\nasync def text_similarity(a: str, b: str) -> float: ...\nasync def text_entities(a: str) -> list[Entity]: ...\nasync def text_tokens(a: str) -> list[Token]: ...\nasync def call_llm(prompt: str) -> str: ...\nasync def call_llm_typed[M](prompt: str, model: type[M]) -> M: ...\n```\n\nthe sidecar writes the scorer once per image, and the sandbox sends back anything it can't run.\n\ncompilerhow much does a new message need the cell division bar chart? write score_message to answer from 0 to 100.\n\n``` php\nasync def score_message(message: str) -> int:\n    text = message.lower()\n    ref = \"cell division bar chart mitotic index data\"\n    sim, nli = await gather(\n        text_similarity(message, ref),\n        nli_scores(message, ref),\n    )\n    score = int((sim * 0.5 + nli[0] * 0.5) * 100)\n\n    def bonus(raw: int) -> int:\n        if \"chart\" in text or \"graph\" in text or \"data\" in text:\n            return min(100, raw + 15)\n        return raw\n\n    return bonus(score)\n    if \"chart\" in text or \"graph\" in text or \"data\" in text:\n        score = min(100, score + 15)\n    return score\n```\n\n1. sandboxrejects draft 1 because nested def bonus() is not allowedrejected\n2. sidecarinlines the bonus, 11 lines instead of 15\n3. sandboxsample message \"which phase dominates in the division chart?\" scores 71 of 100, threshold 65include\n\n- helper call\n- removed from draft 1\n- added in draft 2\n\nWith just these six small helpers, we end up with relatively effective scoring functions (this is a real one!) that are cheap enough to run on every new message.\n\n``` php\nasync def score_message(message: str) -> int:\n    text = message.lower()\n    ref = \"cell division bar chart mitotic index data\"\n    sim, nli = await gather(\n        text_similarity(message, ref),\n        nli_scores(message, ref),\n    )\n    score = int((sim * 0.5 + nli[0] * 0.5) * 100)\n    if \"chart\" in text or \"graph\" in text or \"data\" in text:\n        score = min(100, score + 15)\n    return score\n```\n\nthe `score_message` above, run on five messages. a message that scores 65 or more gets the chart.\n\n`score_message` on the selected message\n\n`async def score_message(message: str) -> int:    text = message.lower()    ref = \"cell division bar chart mitotic index data\"    sim, nli = await gather( text_similarity(message, ref),# sim = 0.625 nli_scores(message, ref),# nli[0] = 0.5    )    score = int((sim * 0.5 + nli[0] * 0.5) * 100)# score = 56 if \"chart\" in text or \"graph\" in text or \"data\" in text:# True        score = min(100, score + 15)# score = 71 return score# 71include = await score_message(message) >= 65include`\n- values for the selected message\n\nsimilarity 0.625 and entailment 0.5 average out to 56, and \"chart\" adds 15, so 71 clears 65.\n\nIt’s then just a matter of asynchronously generating one of these every time an image is sent in a conversation (`history x image -> message -> int`), keeping a [Monty](https://github.com/pydantic/monty) interpreter ready, et voilà! Using this approach has let us be much more free with the number of images we introduce in a long research thread. Give your LLM more DSLs!\n\nthe same ten messages, with the generated scorer added.\n\n- right call\n- in when answering the message needs the chart\n- heuristic\n- in when the message or one of the 2 before it says \"chart\"\n- sidecar\n- a small llm decides on every turn\n- scorer\n- a function the sidecar wrote decides on every turn\n\n- wrong call\n- still deciding\n\neach sidecar call holds its message for about 4,100 ms, shown here 5x faster.the scorer misses only turn 6, at about 20 ms a message against the sidecar's 4,100 ms.", "url": "https://wpnews.pro/news/give-your-llm-more-dsls", "canonical_source": "https://yasyf.com/writing/give-your-llm-more-dsls/", "published_at": "2026-10-05 20:24:58+00:00", "updated_at": "2026-10-05 20:50:21.746861+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "developer-tools", "ai-tools"], "entities": ["Claude Code", "captain-hook", "Aneta", "DSPy"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/give-your-llm-more-dsls", "markdown": "https://wpnews.pro/news/give-your-llm-more-dsls.md", "text": "https://wpnews.pro/news/give-your-llm-more-dsls.txt", "jsonld": "https://wpnews.pro/news/give-your-llm-more-dsls.jsonld"}}