Thanks @John6666 — this is the most useful review the project has received so far, and it landed while the lab was mid-experiment. I’ve adopted the verification/validation split as an orthogonal axis; below is what changed since the original post, where it sits in your chain, and what we’d claim differently now. On the measurement chain Your intervention → render → audibility → named percept → preference decomposition maps cleanly onto what we’re learning the hard way. Our evidence grades (verified-live / verified-doc / source-claimed / hypothesis) turn out to answer a different question than your chain — they say how we know, not which boundary was crossed. We’re now treating them as orthogonal: a mutation can be verified-live on the write path while its musical consequence is still candidate on the percept path. Your failure-state vocabulary (render moved, percept did not etc.) is going into our falsification log format. What’s new since the post
The working repository is now public — codeberg.org/sookoothaii/fruity-project. ~1,000 files: the SONG.py score contract, the tool pipeline, 40+ evidence-graded research docs, and ~227 agent-composed song drafts as MIDI. .flp binaries stay local for now — the MIDI layer demonstrates what the system produces without publishing the sound-design internals.
Offline .flp synthesis works.
offline_bake.py (flpkit-based) bakes a SONG.py into a real FL Studio 2026 project — notes, playlist clips, tempo — without a running DAW, verified by binary readback. FL Studio is now only needed for the final listening judgment, not for producing artifacts.
Two agent instances produced ~200+ compositions across ~80 genres — one on a breadth axis (canonical genres), one on an experimental axis (odd meters, polyrhythm, tempo modulation, form experiments like Basinski-style loop erosion and Reich-style phase processes implemented as parameterized section chains).
A DJ-layer exists: tools/megamix*.py assembles the library into continuous mixes (blend/cut/ambience transitions, tempo-riding incoming material into the outgoing tempo domain — the software equivalent of a pitch fader). Three sets of 40–70 minutes are in the listening queue.
First falsifiable score-level result — a negative one We ran per-song score statistics (drum density, velocity mean/std, syncopation, tension share, pattern diversity, 8-bar energy windows) across the library plus the four operator-rated references (three rated 8/10, one rated 4–5/10 by the only listener who matters here). Aggregate metrics failed to separate the rated-good from the rated-poor track — density, velocity spread, and offbeat ratios were nearly identical. The only candidate differentiator was temporal shape: all three 8/10 tracks show a dip followed by re-escalation past the first peak; the 4–5/10 track decays after its dip. In your vocabulary: this sits at candidate, hasn’t crossed audibility, n=4 rated items — we report it as a hypothesis, not a finding. Failures worth reporting
One agent draft shipped a velocity ramp exceeding 127 (data byte must be in range 0..127) — caught by the render gate, not by ears.
Another shipped an undefined-variable NameError in its score file.
An agent self-audit found and corrected its own metric bugs (a LOO classifier counting positives instead of correct predictions — true accuracy 81.7%, not the claimed figure).
The composition format has a version cliff: five early drafts use a v1 score layout the current renderer can’t read — they’re preserved as unrenderable rather than silently converted.
We had two encoding incidents in the listening index (UTF-8/CP1252 double-encoding); one earlier “corruption” report turned out to be a false positive caused by the IDE rendering UTF-8 as CP1252 — a lesson in verifying bytes before repairing.
Where your chain changes what we do next
Marginal-matched IID vs. structured timing is the right adversarial design for our humanization claims — our GMD-derived profiles currently sit at “matches measured corpus structure”, nothing more.
Relative-part geometry (grid-cancelling same-step timing deltas) is implementable in our humanizer today — it removes the per-hit independence assumption we currently make.
The corpus-dependence warning applies directly: our E-GMD-derived priors describe played drumming; we already tag programmed-vs-played populations, but your point sharpens it — a DnB producer’s grid isn’t a drummer’s limb.
Honest current limits, restated: Windows-only, file-RPC latency, n=1 listener for preference judgments (the operator), most compositions await his ratings, and the re-escalation hypothesis needs more rated data before it earns a stronger state. {ok: true} is still not verification. Repo: sookoothaii/fruity-project: LLM-driven FL Studio production research. SONG.py score contract, offline .flp tooling (flpkit), DJ megamix pipeline, and MIDI draft artifacts. Code + docs + MIDI only; .flp binaries are not published yet. Research-grade, work in progress. - Codeberg.org — the falsification log lives in
research, the listening queue with all MIDI artifacts in
HOEREN.