# YuE2 · Frontier Music with Symbolic Planning

> Source: <https://map-yue2.github.io/>
> Published: 2026-09-11 00:33:26+00:00

Selected score

## From score to song

Listen to a song, then explore the melody, rhythm, and chords in its symbolic plan.

## All selected scores

## Cover & Editing

A familiar song can take a different shape. Listen to changes in melody, lyrics, tempo, and arrangement.

### Agentic music editing

A song takes shape through a conversation. Loading the editing story…

## Genre Explorer

The listening selection, gathered across genres and languages.

## Model & Results

YuE2 (best-of-8) reaches **6.9632** on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench (192 prompts). Suno v5 scores 6.8721 in the same comparison.

## Explore benchmark scores WildSongBench · 15 settings · 7 metrics

**WildSongBench** 192 prompts

| WildSongBench · SongBench |  |  |  | 
|---|---|---|---|
| Rank | System / setting | Prompts | SongBench ↑ | 
|---|---|---|---|

September 5, 2026 evaluation · rankings vary by metric.

[Download results](data/benchmark-results.csv?v=20260908-wsb6)

## How to read these results

**WildSongBench.** 192 prompts and 15 system settings. The table reports automatic evaluation scores. Best-of-8 selects one of eight generations by musicality, prompt control, and lyric accuracy.

**Figure 1.** Song quality combines SongBench and SongEval; text alignment combines MuLan, AllMusicCaps, and prompt control. Both axes show normalized comparison indices. Bubble area represents AudioBox production quality.

#### State of the art on MARBLE

MERT2-30s and MERT2-FS (full-song) achieve **SOTA on 14 of 15 MARBLE metrics**, leading across tagging, key, genre, and emotion recognition.

- SOTA metricsMERT2-30s & MERT2-FS
- 14 / 15
- Genre accuracy · GTZANMERT2-30s · score × 100
- 91.72
- Key refined accuracy · GiantStepsMERT2-FS · score × 100
- 67.05

## Explore MERT2 benchmark scores MARBLE · 15 metrics · 2 encoders

| Scores × 100 · higher is better · bold marks the highest displayed score. |  |  |  | 
|---|---|---|---|
| Benchmark / metric | Best published baseline | MERT2-30s | MERT2-FS | 
|---|---|---|---|
| MTT · TaggingROC-AUC | 91.70PupuJEPA-Large | **91.91** | 91.74 | 
| MTT · TaggingAverage precision | 40.80PupuJEPA-Large | **41.29** | 41.20 | 
| GiantSteps · KeyRefined accuracy | 66.10PupuJEPA-Large | 66.97 | **67.05** | 
| GTZAN · GenreAccuracy | 86.90PupuJEPA-Large | **91.72** | 90.69 | 
| GTZAN · BeatF1 | **91.00** PupuJEPA-Large | 90.59 | 90.57 | 
| EmoMusic · ValenceR² | 62.50PupuJEPA-Large | 63.23 | **63.52** | 
| EmoMusic · ArousalR² | 78.50PupuJEPA-Huge | **80.01** | 78.14 | 
| MTG-Jamendo · InstrumentROC-AUC | 78.40PupuJEPA-Large | **80.27** | **80.27** | 
| MTG-Jamendo · InstrumentAverage precision | 21.20PupuJEPA-Large | 22.89 | **23.51** | 
| MTG-Jamendo · Mood / themeROC-AUC | 76.20PupuJEPA-Large | **79.44** | 78.74 | 
| MTG-Jamendo · Mood / themeAverage precision | 15.50Dasheng-1.2B | **16.68** | 15.74 | 
| MTG-Jamendo · GenreROC-AUC | 86.30AudioMAE++ | **88.01** | 87.98 | 
| MTG-Jamendo · GenreAverage precision | 20.10PupuJEPA-Large / PupuJEPA-Huge | **21.22** | 20.66 | 
| MTG-Jamendo · Top 50ROC-AUC | 83.10AudioMAE++ / PupuJEPA-Huge | **84.18** | 84.13 | 
| MTG-Jamendo · Top 50Average precision | 31.10AudioMAE++ | **32.17** | 31.62 | 

SOTA counts use the best score across the two MERT2 encoders against the nine published baselines in this comparison. Both encoders have 632M parameters. MERT2-30s uses a 30-second training context; MERT2-FS uses 300 seconds. These are full-context representation benchmarks. MERT2 reports the best observed results across representations selected using test scores; each ROC-AUC / AP pair uses the same representation.

#### Six transcription tasks, one model

SheetSage2 achieves **SOTA on 10 of 13 benchmark metrics** with one model for beat, downbeat, key, chord, structure, and melody transcription.

- SOTA metricsOne model · six transcription tasks
- 10 / 13
- Vocal melody · RWC-PopPitch-class note F1 · score × 100
- 82.51
- Chord recognition · osu2017Maj/min · score × 100
- 90.08

## Explore SheetSage2 benchmark scores 6 tasks · 13 metrics

| Scores × 100 · higher is better · bold marks the highest displayed score. |  |  |  | 
|---|---|---|---|
| Benchmark / task | Metric | Best comparison | SheetSage2 | 
|---|---|---|---|
| GTZANBeat | F1 | **88.75** Beat This! | 85.65 | 
| osu2017Beat | F1 | 91.55Madmom | **92.29** | 
| GTZANDownbeat | F1 | 78.28Beat This! | **79.51** | 
| osu2017Downbeat | F1 | 84.99Beat This! | **91.97** | 
| GiantStepsKey | Weighted score | 74.62Madmom | **77.73** | 
| GTZANKey | Weighted score | 74.43S-KEY | **75.77** | 
| osu2017Chord | Maj/min | 84.59Jiang et al. 2019 | **90.08** | 
| Chords1217Chord | Maj/min | **84.09** ChordFormer | 83.81 | 
| HarmonixSetStructure | Accuracy | 80.03SongFormer | **80.51** | 
| HarmonixSetStructure | Boundary F1 · 0.5 s | **70.63** SongFormer | 67.96 | 
| HarmonixSetStructure | Boundary F1 · 3 s | 79.50SongFormer | **82.86** | 
| RWC-PopMelody | Vocal pitch-class F1 | 62.71SheetSage1 | **82.51** | 
| RWC-PopMelody | Full pitch-class F1 | 64.02SheetSage1 | **75.29** | 

SOTA counts refer to the leading scores against SheetSage1, Madmom, and the task-specific systems in this comparison. Results use one model selected by validation loss. Melody F1 measures pitch-class notes; structure F1 measures section boundaries at the stated tolerance. On Chords1217, ChordFormer uses five-fold cross-validation, while SheetSage2 evaluates one fixed model on all 1,217 tracks.

Full comparison includes SheetSage1, Madmom, and task-specific systems.

[Download all results](data/sheetsage2-benchmark-results.csv)

## Training data

Our models are trained primarily on CC0 music and synthetic data. [Tokenwave.AI](https://www.tokenwave.us/) provides most of our synthetic training data under license. We are committed to the ethical and responsible use of data.

- MERT2
- 700K hours
- SheetSage2
- 28.4K hours
- YuE2
- 346K hours
