{"slug": "yue2-frontier-music-with-symbolic-planning", "title": "YuE2 · Frontier Music with Symbolic Planning", "summary": "YuE2 (best-of-8) scored 6.9632 on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench's 192 prompts, according to a September 5, 2026 evaluation, edging out Suno v5 at 6.8721. The same results report MERT2-30s and MERT2-FS reaching state of the art on 14 of 15 MARBLE metrics, including 91.72 genre accuracy on GTZAN and 67.05 refined key accuracy on GiantSteps, while SheetSage2 hit state of the art on 10 of 13 transcription metrics with a single model for beat, downbeat, key, chord, structure, and melody.", "body_md": "Selected score\n\n## From score to song\n\nListen to a song, then explore the melody, rhythm, and chords in its symbolic plan.\n\n## All selected scores\n\n## Cover & Editing\n\nA familiar song can take a different shape. Listen to changes in melody, lyrics, tempo, and arrangement.\n\n### Agentic music editing\n\nA song takes shape through a conversation. Loading the editing story…\n\n## Genre Explorer\n\nThe listening selection, gathered across genres and languages.\n\n## Model & Results\n\nYuE2 (best-of-8) reaches **6.9632** on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench (192 prompts). Suno v5 scores 6.8721 in the same comparison.\n\n## Explore benchmark scores WildSongBench · 15 settings · 7 metrics\n\n**WildSongBench** 192 prompts\n\n| WildSongBench · SongBench |  |  |  | \n|---|---|---|---|\n| Rank | System / setting | Prompts | SongBench ↑ | \n|---|---|---|---|\n\nSeptember 5, 2026 evaluation · rankings vary by metric.\n\n[Download results](data/benchmark-results.csv?v=20260908-wsb6)\n\n## How to read these results\n\n**WildSongBench.** 192 prompts and 15 system settings. The table reports automatic evaluation scores. Best-of-8 selects one of eight generations by musicality, prompt control, and lyric accuracy.\n\n**Figure 1.** Song quality combines SongBench and SongEval; text alignment combines MuLan, AllMusicCaps, and prompt control. Both axes show normalized comparison indices. Bubble area represents AudioBox production quality.\n\n#### State of the art on MARBLE\n\nMERT2-30s and MERT2-FS (full-song) achieve **SOTA on 14 of 15 MARBLE metrics**, leading across tagging, key, genre, and emotion recognition.\n\n- SOTA metricsMERT2-30s & MERT2-FS\n- 14 / 15\n- Genre accuracy · GTZANMERT2-30s · score × 100\n- 91.72\n- Key refined accuracy · GiantStepsMERT2-FS · score × 100\n- 67.05\n\n## Explore MERT2 benchmark scores MARBLE · 15 metrics · 2 encoders\n\n| Scores × 100 · higher is better · bold marks the highest displayed score. |  |  |  | \n|---|---|---|---|\n| Benchmark / metric | Best published baseline | MERT2-30s | MERT2-FS | \n|---|---|---|---|\n| MTT · TaggingROC-AUC | 91.70PupuJEPA-Large | **91.91** | 91.74 | \n| MTT · TaggingAverage precision | 40.80PupuJEPA-Large | **41.29** | 41.20 | \n| GiantSteps · KeyRefined accuracy | 66.10PupuJEPA-Large | 66.97 | **67.05** | \n| GTZAN · GenreAccuracy | 86.90PupuJEPA-Large | **91.72** | 90.69 | \n| GTZAN · BeatF1 | **91.00** PupuJEPA-Large | 90.59 | 90.57 | \n| EmoMusic · ValenceR² | 62.50PupuJEPA-Large | 63.23 | **63.52** | \n| EmoMusic · ArousalR² | 78.50PupuJEPA-Huge | **80.01** | 78.14 | \n| MTG-Jamendo · InstrumentROC-AUC | 78.40PupuJEPA-Large | **80.27** | **80.27** | \n| MTG-Jamendo · InstrumentAverage precision | 21.20PupuJEPA-Large | 22.89 | **23.51** | \n| MTG-Jamendo · Mood / themeROC-AUC | 76.20PupuJEPA-Large | **79.44** | 78.74 | \n| MTG-Jamendo · Mood / themeAverage precision | 15.50Dasheng-1.2B | **16.68** | 15.74 | \n| MTG-Jamendo · GenreROC-AUC | 86.30AudioMAE++ | **88.01** | 87.98 | \n| MTG-Jamendo · GenreAverage precision | 20.10PupuJEPA-Large / PupuJEPA-Huge | **21.22** | 20.66 | \n| MTG-Jamendo · Top 50ROC-AUC | 83.10AudioMAE++ / PupuJEPA-Huge | **84.18** | 84.13 | \n| MTG-Jamendo · Top 50Average precision | 31.10AudioMAE++ | **32.17** | 31.62 | \n\nSOTA counts use the best score across the two MERT2 encoders against the nine published baselines in this comparison. Both encoders have 632M parameters. MERT2-30s uses a 30-second training context; MERT2-FS uses 300 seconds. These are full-context representation benchmarks. MERT2 reports the best observed results across representations selected using test scores; each ROC-AUC / AP pair uses the same representation.\n\n#### Six transcription tasks, one model\n\nSheetSage2 achieves **SOTA on 10 of 13 benchmark metrics** with one model for beat, downbeat, key, chord, structure, and melody transcription.\n\n- SOTA metricsOne model · six transcription tasks\n- 10 / 13\n- Vocal melody · RWC-PopPitch-class note F1 · score × 100\n- 82.51\n- Chord recognition · osu2017Maj/min · score × 100\n- 90.08\n\n## Explore SheetSage2 benchmark scores 6 tasks · 13 metrics\n\n| Scores × 100 · higher is better · bold marks the highest displayed score. |  |  |  | \n|---|---|---|---|\n| Benchmark / task | Metric | Best comparison | SheetSage2 | \n|---|---|---|---|\n| GTZANBeat | F1 | **88.75** Beat This! | 85.65 | \n| osu2017Beat | F1 | 91.55Madmom | **92.29** | \n| GTZANDownbeat | F1 | 78.28Beat This! | **79.51** | \n| osu2017Downbeat | F1 | 84.99Beat This! | **91.97** | \n| GiantStepsKey | Weighted score | 74.62Madmom | **77.73** | \n| GTZANKey | Weighted score | 74.43S-KEY | **75.77** | \n| osu2017Chord | Maj/min | 84.59Jiang et al. 2019 | **90.08** | \n| Chords1217Chord | Maj/min | **84.09** ChordFormer | 83.81 | \n| HarmonixSetStructure | Accuracy | 80.03SongFormer | **80.51** | \n| HarmonixSetStructure | Boundary F1 · 0.5 s | **70.63** SongFormer | 67.96 | \n| HarmonixSetStructure | Boundary F1 · 3 s | 79.50SongFormer | **82.86** | \n| RWC-PopMelody | Vocal pitch-class F1 | 62.71SheetSage1 | **82.51** | \n| RWC-PopMelody | Full pitch-class F1 | 64.02SheetSage1 | **75.29** | \n\nSOTA counts refer to the leading scores against SheetSage1, Madmom, and the task-specific systems in this comparison. Results use one model selected by validation loss. Melody F1 measures pitch-class notes; structure F1 measures section boundaries at the stated tolerance. On Chords1217, ChordFormer uses five-fold cross-validation, while SheetSage2 evaluates one fixed model on all 1,217 tracks.\n\nFull comparison includes SheetSage1, Madmom, and task-specific systems.\n\n[Download all results](data/sheetsage2-benchmark-results.csv)\n\n## Training data\n\nOur models are trained primarily on CC0 music and synthetic data. [Tokenwave.AI](https://www.tokenwave.us/) provides most of our synthetic training data under license. We are committed to the ethical and responsible use of data.\n\n- MERT2\n- 700K hours\n- SheetSage2\n- 28.4K hours\n- YuE2\n- 346K hours", "url": "https://wpnews.pro/news/yue2-frontier-music-with-symbolic-planning", "canonical_source": "https://map-yue2.github.io/", "published_at": "2026-09-11 00:33:26+00:00", "updated_at": "2026-09-11 00:53:09.144803+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-research"], "entities": ["YuE2", "SongBench", "WildSongBench", "Suno v5", "MERT2-30s", "MERT2-FS", "MARBLE", "SheetSage2"], "alternates": {"html": "https://wpnews.pro/news/yue2-frontier-music-with-symbolic-planning", "markdown": "https://wpnews.pro/news/yue2-frontier-music-with-symbolic-planning.md", "text": "https://wpnews.pro/news/yue2-frontier-music-with-symbolic-planning.txt", "jsonld": "https://wpnews.pro/news/yue2-frontier-music-with-symbolic-planning.jsonld"}}