Give your LLM more DSLs A blog post argues that LLMs should be given domain-specific languages (DSLs) to edit rather than free-form prompts, citing a captain-hook example where Claude Code hooks in nested JSON are replaced by a few lines of DSL such as block_command(["rm"], reason="rm is banned in this repo", hint="use trash instead"). The post also describes Aneta's context compiler, which represents data, prompts and examples in a unified IR, runs cache-aware optimization passes, and renders provider-specific output; for one request, no passes sent 9,464 tokens with 8,564 read from cache, COST sent 8,804 with 8,564 cached, QUALITY sent 6,566 with 4,980 cached, and SPEED sent 2,386 with 680 cached. Give your LLM more DSLs + Ask once, compile the answer, run many times It’s been thoroughly shown at this point that having LLMs author prompts for themselves free-form that is, not DSPy-style is disastrous, and that LLM-authored skills actually make things worse https://arxiv.org/html/2602.12670v4 , but writing code is what these next-token-predictors excel at https://arxiv.org/abs/2402.01030 There’s a frequently-used trick for getting LLMs to do out-of-band tasks, which is: come up with a thoughtful DSL, capture your current state as a DSL, get the LLM to manipulate the DSL, then translate that back to your domain. It turns out this works pretty well for in -domain tasks too. A great example of this is the continuous improvement loop in captain-hook https://yasyf.com/writing/less-prompts-more-guardrails/ : getting the LLM to scan transcripts for corrections the user made, then tracking and codifying those using a DSL over Claude Code hooks gives you a nice anti-slop flywheel. your hooks live in nested json, but the model only ever edits a few lines of dsl. { "hooks": { "PreToolUse": { "matcher": "Bash", "hooks": { "type": "command", "if": "Bash rm ", "command": "echo 'BLOCKED: rm is banned in this repo. use trash instead.' &2; exit 2" }, { "type": "command", "if": "Bash pnpm install ", "command": "echo 'BLOCKED: pnpm install at the root. cd into the workspace first.' &2; exit 2" } } } } block command "rm" , reason="rm is banned in this repo", hint="use trash instead", block command "pnpm", "install" , reason="pnpm install at the root", hint="cd into the workspace first", - added Another great example came up recently while building Aneta’s context compiler. The idea behind the context compiler is pretty simple: 1. Take a bunch of data, prompts, and examples and represent them in a unified IR. 2. Run cache-aware optimization passes over that IR every pass weighs the tradeoffs of busting the cache vs decreasing tokens, depending on what that particular call is optimizing for . 3. Render that IR to the provider-specific format for the LLM being called. The trick is capturing a closure at IR-construction time that knows how to render a given piece of context in both verbose and compact ways, and then allowing the optimization passes to toggle one or the other in addition to many standard compaction techniques . The compiler can be optimizing for cost, quality, speed, or some linear combination of those. the same context compiled for COST , QUALITY , and SPEED , where each pass weighs the tokens it saves against the cached tokens it gives up. context nodes, in the order they are senttokens open a node to see both of its renderings. - cache ends here - verbose, 900 tokens php query table arm="b" - {"arm": "b", "n": 12, "index": 0.44, "sd": 0.03, "p": 0.01, "rows": 12 rows of raw counts } compact, 240 tokenssent {"arm": "b", "n": 12, "index": 0.44, "sd": 0.03, "p": 0.01} - read from cache - sent fresh - cut by a pass passes for COST - runscompact the new tool resultsaves 660 tokens, loses none from cache - skipscompact the examples and the tablesaves 3,520 tokens, loses 7,884 from cache - skipsswap figures 1 and 2 for their 24-token captions, since this turn needs only figure 3saves 2,898 tokens, loses 3,584 from cache COST only runs a pass that leaves the cache alone, since a cached token bills at a fraction of a fresh one. | tokens in this request | | | |---|---|---| | objective | total | read from cache | |---|---|---| | no passes | 9,464 | 8,564 | | COST | 8,804 | 8,564 | | QUALITY | 6,566 | 4,980 | | SPEED | 2,386 | 680 | COST sends 8,804 tokens, 8,564 of them read from cache. One of the more interesting optimization passes we do is around images. Images take up HUGE numbers of tokens, yet are often necessary to include in scientific queries as they are figures that are referenced, and OCRing them just doesn’t give sufficient fidelity. However blindly including them results in much larger contexts, and the quality degradation that comes with those. When operating in SPEED or QUALITY modes, the compiler has a pass that allows it to selectively exclude images from the context when compiling along with a tool for the LLM to call if it then decides it needs that image . Except, how do you decide when to include the image? what re-sending every figure on every turn costs, against sending only the one each turn needs. one figure, cut into 28 px patches 1000 x 1000 px - 36 x 36 patches - 1,296 tokens over a thread image tokens sent over 9 turns - re-send every figure - 23,328 tokens - send only the needed one - 11,664 tokens re-sending every figure costs 2x. You can start with simple heuristics like how long ago the image was last mentioned, but that breaks down pretty fast in deep research. Another idea is to have a small sidecar LLM Haiku-class or smaller make the decision at every turn; this works surprisingly well when in QUALITY mode, but is prohibitively slow for SPEED . Inspecting the decisions that sidecar is making yields an interesting insight however: given a static summary of the conversation the decision is a pure function of the message about to be sent. should the chart be included with each message? the right call first, then what each strategy did. - right call - in when answering the message needs the chart - heuristic - in when the message or one of the 2 before it says "chart" - sidecar - a small llm decides on every turn - wrong call - still deciding each sidecar call holds its message for about 4,100 ms, shown here 5x faster.the heuristic answers in under 1 ms and gets 3 of 10 wrong. the sidecar gets all 10 right, but each message waits about 4,100 ms for it. So the question we want to quickly answer is: given a somewhat-recent summary of the conversation, and the message about to be sent, should we include this image? Most of the features the sidecar LLM gathers to make that decision end up being either those easily derived from classic NLP toolboxes, or small binary decisions that an on-device model could make. So how do we handle image inclusion in SPEED mode? You can probably guess where I’m going with this… a DSL It turns out you can just ask the same sidecar model to write a small function, and provide it with a few helpers. Here are the ones we provide: php async def nli scores a: str, b: str - tuple float, float, float : ... async def text similarity a: str, b: str - float: ... async def text entities a: str - list Entity : ... async def text tokens a: str - list Token : ... async def call llm prompt: str - str: ... async def call llm typed M prompt: str, model: type M - M: ... the sidecar writes the scorer once per image, and the sandbox sends back anything it can't run. compilerhow much does a new message need the cell division bar chart? write score message to answer from 0 to 100. php async def score message message: str - int: text = message.lower ref = "cell division bar chart mitotic index data" sim, nli = await gather text similarity message, ref , nli scores message, ref , score = int sim 0.5 + nli 0 0.5 100 def bonus raw: int - int: if "chart" in text or "graph" in text or "data" in text: return min 100, raw + 15 return raw return bonus score if "chart" in text or "graph" in text or "data" in text: score = min 100, score + 15 return score 1. sandboxrejects draft 1 because nested def bonus is not allowedrejected 2. sidecarinlines the bonus, 11 lines instead of 15 3. sandboxsample message "which phase dominates in the division chart?" scores 71 of 100, threshold 65include - helper call - removed from draft 1 - added in draft 2 With just these six small helpers, we end up with relatively effective scoring functions this is a real one that are cheap enough to run on every new message. php async def score message message: str - int: text = message.lower ref = "cell division bar chart mitotic index data" sim, nli = await gather text similarity message, ref , nli scores message, ref , score = int sim 0.5 + nli 0 0.5 100 if "chart" in text or "graph" in text or "data" in text: score = min 100, score + 15 return score the score message above, run on five messages. a message that scores 65 or more gets the chart. score message on the selected message async def score message message: str - int: text = message.lower ref = "cell division bar chart mitotic index data" sim, nli = await gather text similarity message, ref , sim = 0.625 nli scores message, ref , nli 0 = 0.5 score = int sim 0.5 + nli 0 0.5 100 score = 56 if "chart" in text or "graph" in text or "data" in text: True score = min 100, score + 15 score = 71 return score 71include = await score message message = 65include - values for the selected message similarity 0.625 and entailment 0.5 average out to 56, and "chart" adds 15, so 71 clears 65. It’s then just a matter of asynchronously generating one of these every time an image is sent in a conversation history x image - message - int , keeping a Monty https://github.com/pydantic/monty interpreter ready, et voilà Using this approach has let us be much more free with the number of images we introduce in a long research thread. Give your LLM more DSLs the same ten messages, with the generated scorer added. - right call - in when answering the message needs the chart - heuristic - in when the message or one of the 2 before it says "chart" - sidecar - a small llm decides on every turn - scorer - a function the sidecar wrote decides on every turn - wrong call - still deciding each sidecar call holds its message for about 4,100 ms, shown here 5x faster.the scorer misses only turn 6, at about 20 ms a message against the sidecar's 4,100 ms.