cd /news/machine-learning/learning-jazz-pianist-style-with-cro… · home › topics › machine-learning › article
[ARTICLE · art-145715] src=almostimplemented.github.io ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Learning Jazz Pianist Style with Cross-Attention Conditioning

Researchers fine-tuned Aria, a 16-layer transformer pretrained on piano MIDI, on solo performances by twelve jazz pianists from the PiJAMA dataset, adding a gated cross-attention layer that reads a learned four-vector embedding per pianist, according to a paper presented at ISMIR 2026 in Abu Dhabi. A pianist classifier attributed conditioned continuations to the intended pianist 70% of the time versus 37% without conditioning, and a second classifier trained only on generated music identified real recordings with 95% accuracy. The work tests whether a model can compose in a pianist's manner rather than merely recognize the performer, drawing on Dick Hyman's 1994 étude collection In the Styles of… The Great Jazz Pianists.

read5 min views2 publishedOct 5, 2026
Learning Jazz Pianist Style with Cross-Attention Conditioning
Image: source

ISMIR 2026, Abu Dhabi

    In 1994 [Dick Hyman](https://en.wikipedia.org/wiki/Dick_Hyman) published
    [*In
    the Styles of… The Great Jazz Pianists*](https://web.archive.org/web/20160619201634/http://dickhyman.com/Folios/Etudes.htm): fifteen original études, each written in the manner of
    one master, from Scott Joplin to Bill Evans. Rather than transcribing their
    solos, Hyman composed new music that carries their signatures —
    Tatum’s “rapid runs in both hands,” Garner’s
    “strumming, guitar-like left hand,” Peterson’s
    “tremolos and glissandi.” That book is the inspiration for this
    project. Can a model learn to do what Hyman did: not just recognize who is
    playing, but play in their manner? Tatum, Garner, and Peterson are among
    the twelve pianists we study — and so is Hyman himself.
  

    We fine-tune [Aria](https://arxiv.org/abs/2506.23869), a transformer pretrained on piano MIDI, on solo
    performances by twelve jazz pianists from the [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162) dataset, adding a gated cross-attention
    layer that reads a learned embedding for each pianist. To check whether
    the style comes through, we slide a pianist classifier along the
    generated music: conditioned continuations are attributed to the intended
    pianist 70% of the time, against 37% without conditioning. A second
    classifier trained *only* on generated music then identifies real
    recordings with 95% accuracy.
  

Listen first; [how it works](#how) is further down.

The opening bars of “Ain’t Misbehavin’” are played by one of us (Drew). Everything after the dashed line is generated: twelve takes of the same opening, each conditioned on a different pianist. Pick a pianist to hear their take from the top.

Every take generates the same number of notes, so they end at different times: Erroll Garner packs them into 1:26, Cedar Walton spreads them across 2:45. That difference in density is itself part of a pianist’s signature.

…

A shared prompt pulls every pianist toward the same tune. Here each
pianist instead continues a few bars of *their own* playing. The
strip under each take shows what our classifier heard as it slid along
    the continuation, one cell per window of about 300 notes: gold where it
    named the intended pianist, mauve where it named someone else. The two takes per pianist are the best of
    eight we scored; the line under them says how the rest did. Or switch on
    the blindfold and guess for yourself.

Blindfold test: who is playing?

The classifier can also point at moments. On a real performance it is
near-certain almost everywhere, so instead we ask where it is *even
    more sure than usual*: its margin for the true pianist over the
    runner-up, compared with its own average across that performance. Below,
    for one held-out recording per pianist, that curve over the whole piece
    and fifteen seconds from its highest and lowest points.

Click or drag inside a shaded region, or on either piano roll, to play from that point.

    We start from [Aria](https://arxiv.org/abs/2506.23869) (Bradshaw et al., ISMIR 2025;
    [code](https://github.com/EleutherAI/aria)), a 16-layer transformer pretrained on a large corpus of
    piano MIDI. Into each of its last eight layers we insert a cross-attention
    block: the music attends to a small learned embedding for the chosen
    pianist, four vectors per pianist. A learned gate scales what the block
    adds, starting at 0.1, so fine-tuning begins from Aria’s own behaviour
    and learns how much to listen. Because the embedding is attended to at
    every step, the conditioning does not fade as generation goes on, the way a
    prompt prefix does.
  

    How can we tell whether the model has learned a pianist’s style? The
    standard yardstick for a generative model, perplexity on held-out music,
    turns out to be nearly blind to it: given the real preceding notes, the
    next one is predictable whoever is playing, so conditioning barely moves
    the score. Style shows up when the model generates freely and has to stay
    in character on its own output. So instead we let it play, and ask a
    pianist classifier who it sounds like. *Agreement* is how often the
    classifier names the intended pianist, in windows slid along each
    continuation.
Model Perplexity Agreement
Pretrained Aria 11.41 25%
Fine-tuned, no conditioning 6.96 37%
Fine-tuned with pianist conditioning 6.82 70%

Perplexity (lower is better) barely separates the two fine-tuned models; agreement nearly doubles. Continuations are 4096 tokens from 256-token prompts; chance agreement is 8%.

[The paper](paper.pdf) has the details: per-pianist results, the mismatch experiment
    (prompting with one pianist and conditioning on another), memorization
    checks, and a from-scratch classifier that confirms the transfer result.

Everything here is MIDI, rendered in your browser on a sampled piano. The model was trained on automatic transcriptions of commercial recordings, so dynamics and pedalling are approximate, and the rendering is plainer than the records.

The twelve pianists were chosen for separability: they are the

    twelve of [PiJAMA](https://transactions.ismir.net/articles/10.5334/tismir.162)’s thirty whose recordings a pretrained model already
    tells apart most easily. Within them the model imitates some far better
    than others — across the paper’s evaluation, continuations
    were attributed to the intended pianist 96% of the time for Hank Jones and
    Dick Hyman, but only 29% for Cedar Walton.

The scored takes are selected, not copied. They are samples from the corpus of generated music that the paper’s

    synthetic-only classifier learned from; for each pianist we show the two
    highest-scoring of eight candidates. Their prompts come from the training
    recordings, and the paper checks that the continuations do not copy them:
    they resemble their closest training performance less than real held-out
    performances do.

The scores come from a classifier, not from listeners. It is a

    strong one (98.8% of held-out songs), but it has habits: many of its
    mistakes on generated music land on Dick Hyman — fitting, perhaps,
    for a pianist who made a career of playing in everyone else’s style. A listening study is the natural next step.
── more in #machine-learning 4 stories · sorted by recency
── more on @aria 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/learning-jazz-pianis…] indexed:0 read:5min 2026-10-05 · —