# An open research lab for measuring — not prompting — music production in a DAW

> Source: <https://discuss.huggingface.co/t/an-open-research-lab-for-measuring-not-prompting-music-production-in-a-daw/180782#post_3>
> Published: 2026-10-01 02:37:38+00:00

Thanks [@John6666](https://discuss.huggingface.co/u/john6666) — this is the most useful review the project has received so far, and it landed while the lab was mid-experiment. I’ve adopted the verification/validation split as an orthogonal axis; below is what changed since the original post, where it sits in your chain, and what we’d claim differently now. **On the measurement chain** Your `intervention → render → audibility → named percept → preference` decomposition maps cleanly onto what we’re learning the hard way. Our evidence grades (`verified-live` / `verified-doc` / `source-claimed` / `hypothesis`) turn out to answer a different question than your chain — they say *how we know*, not *which boundary was crossed*. We’re now treating them as orthogonal: a mutation can be `verified-live` on the write path while its musical consequence is still `candidate` on the percept path. Your failure-state vocabulary (`render moved, percept did not` etc.) is going into our falsification log format. **What’s new since the post**

**The working repository is now public** — `codeberg.org/sookoothaii/fruity-project`. ~1,000 files: the SONG.py score contract, the tool pipeline, 40+ evidence-graded research docs, and ~227 agent-composed song drafts as MIDI. `.flp` binaries stay local for now — the MIDI layer demonstrates what the system produces without publishing the sound-design internals.

**Offline** `.flp` **synthesis works.**

offline_bake.py (flpkit-based) bakes a SONG.py into a real FL Studio 2026 project — notes, playlist clips, tempo — without a running DAW, verified by binary readback. FL Studio is now only needed for the final listening judgment, not for producing artifacts.

**Two agent instances produced ~200+ compositions** across ~80 genres — one on a breadth axis (canonical genres), one on an experimental axis (odd meters, polyrhythm, tempo modulation, form experiments like Basinski-style loop erosion and Reich-style phase processes implemented as parameterized section chains).

**A DJ-layer exists**: `tools/megamix*.py` assembles the library into continuous mixes (blend/cut/ambience transitions, tempo-riding incoming material into the outgoing tempo domain — the software equivalent of a pitch fader). Three sets of 40–70 minutes are in the listening queue.

**First falsifiable score-level result — a negative one** We ran per-song score statistics (drum density, velocity mean/std, syncopation, tension share, pattern diversity, 8-bar energy windows) across the library plus the four operator-rated references (three rated 8/10, one rated 4–5/10 by the only listener who matters here). Aggregate metrics **failed to separate** the rated-good from the rated-poor track — density, velocity spread, and offbeat ratios were nearly identical. The only candidate differentiator was *temporal shape*: all three 8/10 tracks show a dip followed by re-escalation past the first peak; the 4–5/10 track decays after its dip. In your vocabulary: this sits at `candidate`, hasn’t crossed audibility, n=4 rated items — we report it as a hypothesis, not a finding. **Failures worth reporting**

One agent draft shipped a velocity ramp exceeding 127 (`data byte must be in range 0..127`) — caught by the render gate, not by ears.

Another shipped an undefined-variable `NameError` in its score file.

An agent self-audit found and corrected its own metric bugs (a LOO classifier counting positives instead of correct predictions — true accuracy 81.7%, not the claimed figure).

The composition format has a version cliff: five early drafts use a v1 score layout the current renderer can’t read — they’re preserved as unrenderable rather than silently converted.

We had two encoding incidents in the listening index (UTF-8/CP1252 double-encoding); one earlier “corruption” report turned out to be a false positive caused by the IDE rendering UTF-8 as CP1252 — a lesson in verifying bytes before repairing.

**Where your chain changes what we do next**

*Marginal-matched IID vs. structured timing* is the right adversarial design for our humanization claims — our GMD-derived profiles currently sit at “matches measured corpus structure”, nothing more.

*Relative-part geometry* (grid-cancelling same-step timing deltas) is implementable in our humanizer today — it removes the per-hit independence assumption we currently make.

The corpus-dependence warning applies directly: our E-GMD-derived priors describe *played* drumming; we already tag programmed-vs-played populations, but your point sharpens it — a DnB producer’s grid isn’t a drummer’s limb.

Honest current limits, restated: Windows-only, file-RPC latency, n=1 listener for preference judgments (the operator), most compositions await his ratings, and the re-escalation hypothesis needs more rated data before it earns a stronger state. `{ok: true}` is still not verification. Repo: [sookoothaii/fruity-project: LLM-driven FL Studio production research. SONG.py score contract, offline .flp tooling (flpkit), DJ megamix pipeline, and MIDI draft artifacts. Code + docs + MIDI only; .flp binaries are not published yet. Research-grade, work in progress. - Codeberg.org](http://codeberg.org/sookoothaii/fruity-project) — the falsification log lives in

research, the listening queue with all MIDI artifacts in

HOEREN.
