An open research lab for measuring — not prompting — music production in a DAW An open research lab for measuring music production in a DAW has published its working repository at codeberg.org/sookoothaii/fruity-project, containing roughly 1,000 files including the SONG.py score contract, 40+ evidence-graded research docs, and about 227 agent-composed song drafts as MIDI. The lab reported that offline_bake.py, built on flpkit, bakes a SONG.py into a real FL Studio 2026 project without a running DAW, verified by binary readback, and that two agent instances produced over 200 compositions across roughly 80 genres. Its first falsifiable score-level result was negative: aggregate per-song statistics failed to separate three operator-rated 8/10 reference tracks from one rated 4-5/10, leaving temporal shape as an unconfirmed hypothesis at n=4 rated items. Thanks @John6666 https://discuss.huggingface.co/u/john6666 — this is the most useful review the project has received so far, and it landed while the lab was mid-experiment. I’ve adopted the verification/validation split as an orthogonal axis; below is what changed since the original post, where it sits in your chain, and what we’d claim differently now. On the measurement chain Your intervention → render → audibility → named percept → preference decomposition maps cleanly onto what we’re learning the hard way. Our evidence grades verified-live / verified-doc / source-claimed / hypothesis turn out to answer a different question than your chain — they say how we know , not which boundary was crossed . We’re now treating them as orthogonal: a mutation can be verified-live on the write path while its musical consequence is still candidate on the percept path. Your failure-state vocabulary render moved, percept did not etc. is going into our falsification log format. What’s new since the post The working repository is now public — codeberg.org/sookoothaii/fruity-project . ~1,000 files: the SONG.py score contract, the tool pipeline, 40+ evidence-graded research docs, and ~227 agent-composed song drafts as MIDI. .flp binaries stay local for now — the MIDI layer demonstrates what the system produces without publishing the sound-design internals. Offline .flp synthesis works. offline bake.py flpkit-based bakes a SONG.py into a real FL Studio 2026 project — notes, playlist clips, tempo — without a running DAW, verified by binary readback. FL Studio is now only needed for the final listening judgment, not for producing artifacts. Two agent instances produced ~200+ compositions across ~80 genres — one on a breadth axis canonical genres , one on an experimental axis odd meters, polyrhythm, tempo modulation, form experiments like Basinski-style loop erosion and Reich-style phase processes implemented as parameterized section chains . A DJ-layer exists : tools/megamix .py assembles the library into continuous mixes blend/cut/ambience transitions, tempo-riding incoming material into the outgoing tempo domain — the software equivalent of a pitch fader . Three sets of 40–70 minutes are in the listening queue. First falsifiable score-level result — a negative one We ran per-song score statistics drum density, velocity mean/std, syncopation, tension share, pattern diversity, 8-bar energy windows across the library plus the four operator-rated references three rated 8/10, one rated 4–5/10 by the only listener who matters here . Aggregate metrics failed to separate the rated-good from the rated-poor track — density, velocity spread, and offbeat ratios were nearly identical. The only candidate differentiator was temporal shape : all three 8/10 tracks show a dip followed by re-escalation past the first peak; the 4–5/10 track decays after its dip. In your vocabulary: this sits at candidate , hasn’t crossed audibility, n=4 rated items — we report it as a hypothesis, not a finding. Failures worth reporting One agent draft shipped a velocity ramp exceeding 127 data byte must be in range 0..127 — caught by the render gate, not by ears. Another shipped an undefined-variable NameError in its score file. An agent self-audit found and corrected its own metric bugs a LOO classifier counting positives instead of correct predictions — true accuracy 81.7%, not the claimed figure . The composition format has a version cliff: five early drafts use a v1 score layout the current renderer can’t read — they’re preserved as unrenderable rather than silently converted. We had two encoding incidents in the listening index UTF-8/CP1252 double-encoding ; one earlier “corruption” report turned out to be a false positive caused by the IDE rendering UTF-8 as CP1252 — a lesson in verifying bytes before repairing. Where your chain changes what we do next Marginal-matched IID vs. structured timing is the right adversarial design for our humanization claims — our GMD-derived profiles currently sit at “matches measured corpus structure”, nothing more. Relative-part geometry grid-cancelling same-step timing deltas is implementable in our humanizer today — it removes the per-hit independence assumption we currently make. The corpus-dependence warning applies directly: our E-GMD-derived priors describe played drumming; we already tag programmed-vs-played populations, but your point sharpens it — a DnB producer’s grid isn’t a drummer’s limb. Honest current limits, restated: Windows-only, file-RPC latency, n=1 listener for preference judgments the operator , most compositions await his ratings, and the re-escalation hypothesis needs more rated data before it earns a stronger state. {ok: true} is still not verification. Repo: sookoothaii/fruity-project: LLM-driven FL Studio production research. SONG.py score contract, offline .flp tooling flpkit , DJ megamix pipeline, and MIDI draft artifacts. Code + docs + MIDI only; .flp binaries are not published yet. Research-grade, work in progress. - Codeberg.org http://codeberg.org/sookoothaii/fruity-project — the falsification log lives in research, the listening queue with all MIDI artifacts in HOEREN.