# Jamming with Jev and Claude on TidalCycles

> Source: <https://jdsemrau.substack.com/p/jamming-with-jev-and-claude-on-tidalcycles>
> Published: 2026-10-05 11:37:40+00:00

#### What is creativity?

When thinking about what makes humans more intelligent than machines is our ability to create outstanding works of art. Our Western Schelling Points are that we immediately start to think about Leonardo Da Vinci’s Mona Lisa, Wolfgang Amadeus Mozart’s Requiem, or Wanderer Above the Sea of Fog by Caspar David Friedrich. Even outstanding achievements of technology like a Ferrari, an Apple laptop, or buildings like the Sagrada Família in Barcelona can be seen as works of art. We sometimes say that these works touch our soul.

But what does that mean?

With AI agents becoming substantially more intelligent, are they also creative?

I got a couple of free credits from Anthropic, and I used them for building a coding agent that, with the help of TidalCycles, can generate infinite music.

Similar to previous research into [vindication](https://medium.com/@jsemrau/game-theory-and-agent-reasoning-i-eed64b49ff37) and [ambition](https://jdsemrau.substack.com/p/governing-ambition-with-owasp-llm), I am now trying to answer the question: if infinite monkeys type on infinite computers, will one of them write Mozart’s Requiem?

Before you spend more time reading my post, here is a result how this sounds like.

#### Table of Contents

1. Creativity as a property of a system
2. What makes a song good in mathematical terms?
3. Live coding as a stage for machine creativity
4. Designing from first principles
5. The architecture
6. The phase machine
7. The director’s instruction
8. The role of surprises
9. The role of the judge
10. The code model
11. Prompting Tidal
12. Results: The live engine
13. Divergence: does it produce novel ideas?
14. Lessons Learned
15. Sources

So let’s start to define first what we understand as creativity.

Psychology’s standard definition needs three things:

1. An **original** and**effective** creative product. Something new that doesn’t work is noise. Something that works but isn’t new is routine.
2. Margaret Boden adds in The Creative Mind a third one: **surprise**

Another minor variation is to distinguish the character of the innovation, whether it is psychologically creative (P-creative) when it is new to the one who creates it, and historically creative (H-creative) when it is new to everyone.

For music, that means each track must be considered new compared with what other music the listener has listened to already.

Boden further spreads creativity along three axes.

1. Combinational creativity that joins existing ideas in unfamiliar ways. The smartphone is a typical example here.
2. Exploratory creativity that moves through a structured space, like a chessboard or a blues rhythm, and finds new ways to play.
3. And finally, transformational creativity that changes the rules of the space itself. Einstein’s Theory of Relativity is the prime example here, because he changed Physics itself.

Most creative innovation today is combinational, as the world around us is constantly improved by research institutes in the world around them, and new researchers are graduating to make their mark on the world.

Transformative creativity is rare. Even among humans.

Following that definition, my “Jam” session could be seen as combinatorial, as it combines music as code, autonomous agents, and live streaming with user input in a new way.

We can argue that the invention and successive productization of generative AI fundamentally transformed the world of Artificial Intelligence. I sometimes refer to it as a completely new branch. Even though there is a clear lineage from Neural Networks and Markov Chains to Large Language Models.

A language model by itself can never be truly creative, as it is trained to predict the most likely continuation. Therefore, by definition, it will always choose the most familiar way in a structured space. Ask the model for a chord progression, and it writes the most typical one.

This is also sometimes referred to as the Exploration vs Exploitation problem.

Instruction tuning narrows its range further, a narrowing often called mode collapse. Raising the sampling temperature doesn’t solve this. It adds randomness, which raises novelty but lowers quality at the same time.

It then generates music that is neither pleasant nor interesting to listen to. And I don’t mean that in a Hindustan way.

#### Creativity as a property of a system

Csikszentmihalyi’s systems model (1988) places creativity not in one mind but in the interaction of three parts:

1. an individual who produces variations,
2. a domain of rules and symbols, and
3. a field of gatekeepers who decide which variations survive.

The Geneplore model (Finke, Ward and Smith, 1992) splits the creative act into generating candidates who structure and exploring them.

The agent system does this with modern agent tools; such a system is apparently obvious.

What does this mean for my project?

Let’s start with a general flow of how such an agent could work.

**Generation**. Here  I have implemented 2 agents. The first is the Director, who sketches the overall track and adjusts execution with verbal instructions per track. The second is the Performer, who edits or generates three proposals per move.

**Domain** contains the art direction. I.e.,  the rules and symbols that are valid for this track, the validator, the harmony rule, the block library, kits, and valid ranges.

**Field** the Jev-based judge in the loop, which classifies the most probable statement, and the human-in-the-loop that gives me a simple 👍/👎 per line, adjusting the song during flight.

**Memory** is an important component in all my projects, and here it’s the same. This time the project contains the track’s history, the last four tracks’ keys, progressions, hooks, and surprises, and the session logs.

**Chance** is where creativity comes in. Here, a planned accident, not a completely random spark, shapes the arc of the song.

Let’s try to model creativity **C** mathematically in the context of generating music. Creativity is expressed as the current plan in relation to the history of making music.

Let **x** be the director’s plan for a track. 

**L** be the space of all valid plans (in key, chord tones on strong steps, sounds from the style’s kit, values in range) and 

**H** be the stream’s memory of the last four tracks. 

Then we can write creativity as

Given the stream's recent history, H is its value times its novelty. Because both factors are 0 or 1, multiplying them works as an **AND**. A plan that is valid but repeats something scores 0, and so does a fresh plan that's out of key. "New" is also only relative to`H`: the plan only has to be new for this stream's recent tracks, not new in music history. `C = 1` is true when the plan is valid, and none of its key, progression, or surprise was used in the last four tracks. Otherwise `C = 0`.

Value is **pass or fail**: 1 if the plan lies in `L`, the set of valid plans (in key, chord tones on strong steps, sounds from the style's kit, values in range), and 0 otherwise. There is no "more valuable" plan. That is what the next sentence in your doc means by "value is a gate, not a goal".

Novelty is checked separately for three features of the plan:

- `f(x)` is the plan’s value for that feature, e.g., its key, “A minor”.
- `H_f` is the set of values that feature had in the last four tracks.
- Each factor is 1 if the plan’s value is not among those four.

The product (`∏`) is again an AND, so the plan is novel only if its key, its chord progression, and its surprise are all different from the last four tracks.

In this concept, value is a gate. But in its actual code implementation, the gate repairs rather than rejects. The review function replaces each illegal part of a plan with a legal one and logs the correction. f(x) is one feature of the plan, its key, for example, and H_f the values that feature took in the last four tracks.

Novelty is enforced for the three features in the product:

1. a key from the last four tracks gets a new root in the same mode, and
2. a recent progression or surprise is replaced, or
3. other features (the hook, the shape of the sections, the sounds) are asked to differ in the prompt but not enforced.

#### What makes a song good in mathematical terms? 

The agent samples its plan from a dynamic list of all possible and valid plans.

Where b is the fixed brief (the vocabulary and rules, the same for every track), sigma is one of twelve sparks drawn at random (”a hook with more rests than notes”, “start the progression away from the tonic”), and hat k is a random suggested key that isn’t in H, half the time a neighbour of the previous key. sigma and at k move the model’s starting point without lowering value, because whatever it writes still passes the gate.

Within a track: complexity as a target. A track is a creative moment-to-moment process that keeps developing.

The complexity of a program P can be expressed as

Each phase of the arc has a target

with

After two moves in a row that only turn a knob, the next move must change the music itself. After running a series of experiments, I landed on the rule to have at most one surprise per track and never one of the last four, chosen by the director from a menu of eight rule-breaks built only from safe operations.

Surprise in Boden’s sense is a violated expectation, and here the expectation is the one the music sets up itself: the drop that falls back, the kick that stays away. If the agent contributes anything beyond the space the code designed is measured with the blind Judge

where pi is the share of blind comparisons the agent’s ideas win against random-but-legal ideas, and rho is the share they would win by chance (their share of the pool, roughly 0.5).

Then delta > 0 means an independent judge prefers the model’s choices to chance within the same space.

Code makes every option good enough, memory makes the chosen option new, chance pushes the model off its most typical answer, the surprise budget lets it break the arc on purpose, and a blind judge checks that it chooses better than chance. In Boden’s terms, the system is P-creative and exploratory, with a small, pre-approved step towards transformation in the surprises.

To give you a more hands-on expression from one sample session with Claude Sonnet 5 as director and performer, Jev as blind judge, trance, about 71 minutes and 12 tracks.

Where novelty was not enforced, the model’s typical answer came back. All 12 progressions start on the tonic chord, even with a spark asking for the opposite. Eleven of the twelve hooks start on the root and outline the triad. Four titles begin with “Iron” and three with “Concrete”, “dark, driving” recurs in the moods, and every tempo falls between 136 and 144 BPM of the 126 to 146 allowed. Asked to change the subgenre from track to track, the model chose psy six times, acid four times, uplifting twice, and progressive never.

The first lesson was that novelty appears where it is enforced or measured, and typicality returns everywhere else.

But I am getting ahead of myself.

For an agent system, creativity is therefore partly an engineering decision: which dimensions to remember, which to perturb, and which to leave to the model.

Since I want to ensure we are not generating the full track as Suno does it, and also want to emphasize the live editing nature, I landed on live coding.

#### Live coding as a stage for machine creativity

Live coding is a niche performance practice in which a human musician writes and modifies code while it plays, with the code projected for the audience.

In TidalCycles, a Haskell-embedded pattern language, a set might start with a single line:

`'d1 $ s “bd ~ ~ bd ~ ~ bd ~'`

From this starting point, the code may grow into a small program of up to nine lines, one per channel ('d1'...`d9') Each line is a pattern: rhythms and melodies in compact mini-notation (‘0 ~ 2 4’, ‘<a b> ’), through effects (‘# lpf 800’, ‘# room 0.4’) and transformations (‘jux rev’, ‘every 4 (fast 2)’, ‘off 025 (|+ n 12)').

Re-evaluating a line replaces that channel’s pattern at the next cycle. The music never stops; it is steered.

What the audience enjoys is not the final program but its trajectory during the live coding session: the build-up, the moment the acid filter opens, the breakdown when the kick drops out, the small twist that makes a loop feel new.

This immediate audible feedback makes live coding an unusually good test of creative behaviour in language models, for four reasons:

1. It is open-ended. No edit is correct; edits are only better or worse in context.
2. It is sequential and self-referential. Every decision must be read against the program as it is now, which is the model’s own earlier work.
3. It is long-horizon. A stream runs for hours. A model that has three good ideas and then loops on them fails visibly.
4. Judged by ear. Listeners, not unit tests, decide what is good. That keeps the test honest about what creativity means.

If you are asking a typical model to write a track in one shot tests something different: recall of syntax seen during training. Local models (for example, quantized 27B to 35B Qwen models served by ‘llama.cpp’) do poorly at this. They produce type errors, wrong notes, clashing layers, and walls of sound.

None of that tells us whether the model has musical judgement.

#### Designing from first principles

Separate the space of legal moves from the choice among them.

The system, not the agent, guarantees that a move is syntactically valid, in key, within sensible ranges, and appropriate for the instrument. The model only chooses among legal moves. Anything its choices add beyond a random walk through the same moves is its creative contribution, and that contribution is measurable.

The artefact is the code. The program is not trained on music and thus does not inherit an internal hidden representation besides the code. Everything shown in the code window, it is what the model reads, and it is what gets revised.

One small edit at a time. Similar to popular live-coding streamers, one edit touch should be small, and this mimics how a human performer would engage with the code. This keeps every decision attributable (so it can be evaluated), and limits the damage a bad decision can do. This also increases continuity and homogeneity.

The music never stops. What I wanted to understand is if there is a point where songs repeat or return garbage. While the performer takes over any edit the model fails to provide, so model failures become data points, not total outages. I noticed that over longer periods, music would become repetitive in earlier versions.

I improved memory to not repeat that.

Understanding *why* each agent made a specific decision is as important as the music. Every LLM exchange is shown on screen as it happens (who asked which model what, and what came back); every edit is tagged with who chose it; the model states what it hears before it states its move, and the system’s own structural decisions (the track’s arc) are logged with their reasons.

#### The Architecture

The program model, in `music/livecode.py` and `music/theory.py`, holds the live program as data and renders it to Tidal. It defines the edit language and the validator that checks each edit, and it holds the keys, the modes, and the progression library.

The track plan, in `music/plan.py` and `music/surprises.py`, holds the Director’s plan for each track (the sketch, the cues and the surprise) along with the checks on that plan. It also holds the menu of surprises.

The block library, in `music/blocks.py` and `tidal/synthdefs.scd`, provides curated instruments and starting lines for each of the nine styles, as well as the trance subgenres and their kits.

The mixer, in `music/mix.py`, sets levels calibrated from measured loudness, handles fades and density ducking, and enforces the gain window.

The agents live in `music/performer.py`, `music/director.py`, `music/live_engine.py` and the `prompts/` folder. They are the director; the performer, with its phase machine, prompt, schema, and built-in fallback; the Judge; and the Narrator.

The live engine, in `music/live_engine.py` and the `Clock` in `music/engine.py`, handles timing. It applies edits on the bar, sends only the changed lines, and publishes the current state.

The LLM gateway, in `music/llm.py` and `audio_stack.py`, routes requests to the right provider (local, Claude or jevtypesafe) and constrains decoding to the schema. It also manages the local server’s lifecycle and timeouts, and logs every exchange.

The audio and stream layer, in `audio_stack.py`, `tidal/`, `stream/` and `music/narrator.py`, runs the GHCi session and boots SuperCollider. It also manages the master bus, the TTS narration, the stream page, ffmpeg, and the watchdog.

One edit cycle looks like this.

While the current bars play (8 bars in slow styles, 16 in club styles, about 20 to 30 seconds), a worker thread builds a prompt from the current program and asks the model for its next move, under a JSON schema of the moves that are legal right now. In judge mode, the model proposes three moves, the built-in performer adds three, and the Judge picks one without knowing who proposed what. On the bar line, the engine validates the chosen edit and sends only the changed Tidal statements: a new line fades in, a removed one fades out, and lines whose level changes because the mix got fuller are re-sent. It then updates the on-screen code, logs what the model heard, what it did and why, saves state, and, if the track’s arc moved on, asks the Narrator for a line. If the reply is late, invalid, or missing, the built-in performer’s edit is used instead, tagged as such.

The loop then starts over.

“The LLM makes music” is too vague to design or evaluate and will likely fail.

So the architecture assigns explicit roles. Each role has a defined input, an output contract, a validator, and a fallback. The live engine uses the first five; the last two belong to the older arranger engine and are described because contrasting them is instructive.

What are the different components and what role does the LLM take?

The Performer is the central role, and the only one that touches the program. The model sees the program exactly as the audience does, plus a machine-readable view of what it may change, and returns **exactly one** operation (three alternative proposals in judge mode, section 9.5). It never writes a Tidal statement itself. It writes *edits*, which the system turns into statements. This is what lets a 35B local model play for hours without a single syntax error reaching the audio engine.

What the Performer decides:

- **Arrangement** : which instrument enters next, when to drop the beat, when to strip back.
- **Composition** : melodic contour and rhythm, as scale-degree and mini-notation strings.
- **Timbre and space** : filters, reverb, delay, the acid knob, including slow sine modulation.
- **Transformation** : which Tidal higher-order functions to wrap around a line.
- **Energy** : small tempo changes.

#### The phase machine 

Unstructured freedom produces aimless music. Every track follows a phase machine that gives it a dramatic arc and tells the model what kind of move fits now:

build ──(target lines reached)──▶ groove ──(6-10 edits)──▶ breakdown ──(4 edits)──▶

rebuild ──(muted lines back)──▶ groove2 ──(6-10 edits)──▶ outro ──(≤1 line left)──▶ next track

The phase machine is organized by the director agent.

#### The director’s instruction

The director, who plans this track, says: “Thin the acid to a slow pulse (the acid line).”

Carry it out: make the **ONE** edit that does what the director says, in code. The fake drop needs the cue queue too: during the false peak, the real peak’s cues wait. A complexity target per phase. Every prompt carries one line that places the music against its target: When a groove is more than 4 below its target, the performer edits twice as often as a live coder does while building a section up. 

The engine reports per track how close it got: complexity in the groove: 13.5 on average, 15.0 at the peak. A model that is unsure only turns knobs: a filter a little lower, the reverb a little wider.

That is safe, and after a few minutes it is boring.

Before each track of about five minutes long, the director writes a plan detailed enough to build the track from, and instructions for every part of it and stores it as one JSON object:

```
{”do”: “start”, “part”: “acid”, “block”: “acid”, “sound”: “ys_acid”,

 “pattern”: “0 4 0 7 0 4 2”, “say”: “Start the acid: rolling seven-step line”}

{”do”: “knob”, “part”: “acid”, “knob”: “acid”, “to”: [0.3, 0.8, 16],

 “say”: “Open the acid over sixteen bars”}

{”do”: “stop”, “part”: “kick”, “say”: “Drop the kick”}
```

The actions are ‘start’, ‘stop’, ‘bring back’, ‘pattern’, ‘nob ’, ’ sound ’, ‘sound ’, ’ wrap ’, and ‘tempo’. At the start of each part of the ac thePerformer carries out that part’s cues first, one per move. A cue that names everything (part, block, sound, pattern, value) is played exactly as given. A cue that leaves something open (”busier, in 6ths, climbing”) goes to the performer, who writes the code for it.

The prompt (‘prompts/director_piece.md’) has two parts.

1. The brief comes first and is the same for every track: the role, the vocabulary(styles, subgenres, chords per mode, block sounds, knobs, voicings, hook parts), what to write, and the reply format, about 2,900 tokens.
2. The context follows: the previous track and its shape, the keys, hooks, and progressions of the last four tracks, he listeners’ dislikes, a suggested key, a spark, and the surprise nu, about 400 tokens. The order is deliberate. Claude caches the brf (section 11.4), and the variable part that comes last keeps the model’s attention on what is new.

Review checks every field and repairs what fails: a key from the last four tracks gets another root in the same mode, a recent progression is replaced, a cue naming a sound the part can’t play is fixed or dropped, a tempo outside the style’s range is replaced.

It works, but it quickly gets boring. In order to give the listerner something new we need surprises.

#### The role of surprises

The table below lists eight pre-defined rule-breaks, each built only from moves the system already makes solely (mute, unmute, the drop, tempo, key, the arc plan), so none can break he harmony or the flow:

The agent needs to decide whether a surprise is allowed (every second track by default, not one of the last four, and only if the track has what it needs: no ‘solo_breakdown’ without a breakdown). The director decides which one suits the track. Code places it at its moment, and the streampage announces it when it fires (section 10.4). The menu is the answer to a trade-off the host named: a model can’t judge how risky an idea is, so it shouldn’t invent risks, but a stream without surprises gets boring over 24 hours. The menu gives it bounded risks to choose from.

Before the live engine, the director was the main role and chose presets per section from catalogues. That proved reliable but a poor test of creativity: decisions were coarse, and the music was correct but static, and listeners called it boring. The live engine first reduced the director to a key, tempo, progression, and mood. It came back as an art director when the music turned out samey and its instructions vague. It now writes material (chords, a hook, patterns) and precise direction, and the Performer still makes every move in between.

#### The role of the judge

When Jev came out, it solved my classification problem, and I started using Jev as an inexpensive decision engine.

The performer evolved into a generator: it proposes now three different moves in one call, the built-in performer adds up to three of its random-but-legal moves, and a decision model (jevtypesafe’s typed decision API) picks one from the shuffled, unlabelled pool.

The Jev-based judge returns a probability per option:

```
💬 Judge ← jevtypesafe/decide (0.9s): {”edit”: {”choice”: “e2”, “confidence”: 0.26,

                                        “probabilities”: {”e2”: 0.38, “e6”: 0.34, “e5”: 0.21, ...}}}

⚖️  Judge chose the LLM idea (p 0.38, confidence 0.26, from 3 LLM + 3 built-in)
```

Because the Jjudge is blind, how often it prefers the model’s ideas over random ones, and how much probability it gives them, is a direct measure of the model’s creative contribution. A decision API cannot write text, so it cannot perform alone in any interesting way: given only built-in candidates, it can only pick the best of random ideas. The pairing is the point: a generator for ideas, a judge for taste.

#### The code model

In music as code, a program is a track (key, tempo, progression) plus up to nine lines, one per musical role. Roles are pinned to Tidal channels so the screen always reads like a hand-written set:

```
ROLE_CHANNELS = {
                  ”beat”: 1, “bass”: 2, “chords”: 3, 
                  “keys”: 4,“lead”: 5, “acid”: 6, 
                  “texture”: 7, “perc”: 8, “arp”: 9
                }
```

Keeping the parts separate is what makes small, meaningful edits possible.

Every edit is one of eleven operations:

The language is deliberately focused. In order to be cost-efficient, the model needs tp learn everything it needs from the prompt, and each operation corresponds to something a human live coder actually does.

#### Prompting Tidal

The prompt (‘prompts/performer.md’) is built fresh for every edit.

Early versions gave the model names as ranges only: it saw 'off *0.25 (|+ n 12)*’ or ‘*lpf 200-9000*’ but not what they sound like, and its edits felt random. 

With these prompting techniques the context became clearer.

1. **Framing** : “You are performing a live-coding set in TidalCycles, like[Switch Angel](https://www.youtube.com/watch?v=iu5rnQkfO6M) ... The audience sees the code and hears every change, and they read your notes to understand why,” plus the stream theme and style.
2. **The program as code** : exactly what viewers see.
3. **Phase** , brief and example moves for this phase: local models copy good examples far more faithfully than they follow rules (”in a breakdown: mute the beat first; the space it leaves is the point”).
4. **Memory** of its own recent edits, the last six that actually played, with reasons: “build on these; don’t undo a move you just made, and don’t make the same kind of move three times in a row”. Without this memory, the model only sees the present and tends to loop or undo itself.
5. **Memory** : A**Tidal cheat sheet** : mini-notation (‘~ [ ] > * ! ?,'), what scale degrees mean, and one line per wrapper and effect on what it sounds like (‘jux rev’: a reversed copy in the other ear, eerie and dreamy; ‘lpf’: low is dark and muffled, high is bright and open). Only the wrappers and effects that apply to the line playing now are explained, which keeps the prompt small.
6. Editable parts as compact **JSON** : per line its slots and values, its effects, which effects it accepts and which wrapper group it takes.
7. **Addable blocks** , operations, taste guidance: aim for the richness of a real live-coded set (up to ‘max_lines’ interlocking lines, things always moving), harmony first, and make each edit something a listener will notice (new notes, a rhythm, a layer, a sweep) rather than a reverb nudge. An earlier version said “prefer revising over adding”
8. **Response** : `listening`, `reason`, the operation, optional `narration`; or, in judge mode, three genuinely different proposals.

#### Results: The live engine

Let’s start with some token cost considerations. Over a 71-minute session, the director read 59,000 tokens from the cache and stored 5,363. Using this, I could reduce its input cost by 34% using prompt caching. The performer, whose program and recent edits change with every move, came to a reduction of 54%.

Claude is too expensive to run 24/7, and the local model’s plans are notably weaker.

The future plan is to fine-tune the local smaller model with Claude’s answers

#### Divergence: does it produce novel ideas?

Three measures show whether the model generates ideas of its own.

Operation entropy is the **Shannon entropy** over ‘op’. A performer that only rides lpf scores near 0, and one that uses all eleven operations evenly scores about 3.46 bits. Novel material counts the distinct pattern values the model writes that appear neither in the block library nor earlier in its own history. Self-repetition is the share of edits that repeat an identical (op, role, name, value) from the previous k edits, or undo one (A→B→A).

**Phase fit** asks whether the performer strips back in the breakdown and adds during the build; a simple rule table scores each edit against its phase brief. Motivic development is the edit distance between successive values of the same melodic slot. Small distances mean development and large ones mean replacement, and a musical performer shows both, mostly the first. '

**Reason consistency** checks that reason matches the edit, so “darken” should lower lpf and “open the acid” should raise acid. Rules cover the common verbs, and an LLM judge or a human covers the rest. Breadth of attention is the distribution of edits over roles, since a performer that edits only the lead is ignoring the ensemble.

**Judge mode** gives a measure that needs no listening panel. For every edit, the model’s ideas compete with random but legal ideas in front of a judge that cannot see where each came from. This gives two numbers per track. The preference rate is how often the judge chose the model’s idea, against a chance level of roughly the model’s share of the pool (about 40-50%). The probability mass is the average share of the judge’s probability given to the model’s ideas, which also captures near misses.

```
[Performer] the judge preferred the LLM's idea 14/20 times (70%) over the built-in
            candidates; on average it gave them 61% of its probability
```

Because the judge has its own taste and biases, this measures agreement with one judge rather than absolute quality. Comparing several models under the same judge, and the same model under several judges, separates the two.

#### Lessons Learned

**Constrain the vocabulary** Scale-degree slots, curated blocks and ranged effects remove whole classes of errors while leaving thousands of distinct musical moves open. Even smaller local models that could not write a valid Tidal track on their own perform sensibly inside this frame.

**Make the code the artefact.** Coarse planning roles like the director produced correct but lifeless music. Once the program itself became the thing the model grows and revises, both the music and the model’s decisions became legible.

**Make decisions frequent, small, and attributable.** Decisions like these sound like performance, and they make evaluation possible.

**Use fallbacks to turn failures into data.** Continuity is what makes long-horizon testing possible at all.

**Ask for reasons.** One line of intent costs a few tokens. It is the single most useful artefact for judging whether a model “meant” what it played.

**Treat the validator as the security boundary.** When model output drives a live interpreter, whitelists are not optional.

**Fix failures at the source.** A local model’s unusable replies were mostly malformed names and slightly-off numbers. Three changes did more than any prompt wording: constraining decoding with a schema of the legal moves, clamping numbers, and allowing one corrected retry.

**Teach the language in the prompt.** A model that only sees `off 0.25 (|+ n 12)` has to guess what it sounds like. Three additions made the edits noticeably more purposeful: one line per wrapper and effect describing its *sound*, example moves for each phase, and the model’s own recent edits.

**Treat explanations as a product, not a debug aid.** The system shows what each agent heard, what it did and why, and tags every edit with who chose it. That is what lets a human trust or distrust the music, and it is the data a creativity evaluation needs.

**Measure the mix instead of guessing it.** Hand-set levels were 4-8 dB off for sustained beds. Measuring every synth offline and deriving gains from targets fixed what listeners complained about. Fades and ducking make an addition sound like part of the song rather than a switch.

**Pair a judge with a generator.** A decision model alone can only pick the best of random ideas. Paired with a generator, it becomes both a source of taste and a blind measure of the generator’s ideas.

**Listen, then measure.** Every major fix started with a listener saying what was wrong (”too loud”, “not in harmony”, “much simpler”). Every one ended with a number that pins the problem down, such as a level target, a chord-tone rule, or the number of lines at the peak. The first diagnosis was often wrong; the ear was not.

**Treat scaffolding as a trade.** Each guardrail that improved the music took a choice away from the model. Track the model’s share of the decisions alongside the quality of the music, or the code quietly becomes the composer.

**Budget the context.** A prompt that grows with every useful addition eventually overflows a small local context window. Describing only what applies now, and falling back to a compact prompt when needed, keeps the performer playing.

**Enforce novelty where it matters.** A model asked to vary something returns to its typical answer unless the variation is checked. Decide which features must never repeat, remember them and check them. Leave the rest to the prompt and measure how far it drifts (section 1.5).

**Give a freedom only with its check.** Letting the Director write its own chords and hooks improved the music where earlier freedoms had made it worse. The difference was that each new freedom came with the rule that keeps it musical.

**Make instructions between agents precise.** A director’s note that reads well (”let it breathe”) gives the next agent nothing to do. A structured cue with exact names can be played, checked, and shown.

**Budget risk with a menu.** A model can’t judge how risky an idea is. A short list of rule-breaks, each built from safe moves, lets it take risks without having to invent them.

**Put stable content first and variable content last.** Ordering every prompt from what never changes to what changes each call lets a local server reuse its cache. It also cuts a hosted model’s input cost by a half to two-thirds, at no cost in quality.

#### Sources

- [TidalCycles documentation](https://tidalcycles.org/docs/)
- [Switch Angel](https://www.youtube.com/@Switch-Angel) ‘s live-coded trance sets and scripts, the musical reference for the trance style
- [NRG-CP chord progression dataset](https://doi.org/10.5281/zenodo.15304989) (WaivOps / Patchbanks), CC BY 4.0
- M. A. Boden, *[The Creative Mind: Myths and Mechanisms](https://www.routledge.com/9780415314534)* (1990; 2nd ed. 2004)
- M. Csikszentmihalyi, “Society, culture, and person: a systems view of creativity”, in R. J. Sternberg (ed.), *The Nature of Creativity* (1988)
- R. A. Finke, T. B. Ward and S. M. Smith, *Creative Cognition: Theory, Research, and Applications* (1992)
- M. A. Runco and G. J. Jaeger, [“The standard definition of creativity”](https://doi.org/10.1080/10400419.2012.650092) ,*Creativity Research Journal* 24(1), 2012
- The 24/7 stream: [@infinitrance on YouTube](https://www.youtube.com/@infinitrance)
