{"slug": "machine-learning-of-artistic-fingerprints-in-jazz", "title": "Machine learning of artistic fingerprints in jazz", "summary": "Researchers trained supervised learning models to identify 20 iconic jazz musicians from 84 hours of recordings, achieving 94% accuracy with a multi-input architecture that separately represents melody, harmony, rhythm, and dynamics. The study, published in Nature Machine Intelligence, releases open-source implementations and a web application for exploring results.", "body_md": "## Abstract\n\nArtists are often recognizable through collections of distinctive patterns (‘fingerprints’) in their work. Identifying such traits has important applications in authorship attribution, education, cultural heritage research and historical analysis. Here we focus on music, a domain with a rich tradition of theoretical and mathematical analysis. We train a variety of supervised learning models to identify 20 iconic jazz musicians from a curated dataset of 84 h of recordings. In particular, we introduce a multi-input architecture that represents four musical domains separately: melody, harmony, rhythm and dynamics. This design allows us to accurately identify individual performers (our best model obtains 94% accuracy across 20 classes) and to examine which musical elements most strongly distinguish between individual artists. We release open-source implementations of our models and an accompanying web application for exploring our results.\n\n### Similar content being viewed by others\n\n## Main\n\nWhat distinguishes one artist from another? This question lies at the heart of arts scholarship. In the visual arts, distinguishing cues might include subject matter, colour palette and brushwork; in literature, vocabulary, syntactic patterns and narrative archetypes; and in music, harmonic progressions, rhythmic structures and melodic motifs. When multiple authorial cues are considered together, they come to make up the distinctive authorial fingerprint of an artist.\n\nSome of these cues are perceptible to humans, even for a relatively untrained eye or ear: for example, the way a painter renders light at the edge of a shadow, or a novelist’s habitual reliance on certain words or sentence structures. With training, humans can learn to identify more subtle authorial cues, such as slight differences in word distributions between two writers<sup>[1](/articles/s42256-026-01279-9#ref-CR1)</sup>. Other artistic cues may lie outside the realm of immediate human perception: for example, the chemical composition of the paint used in an artwork<sup>[2](/articles/s42256-026-01279-9#ref-CR2)</sup>, or the microtiming deviations in a musician’s performance<sup>[3](/articles/s42256-026-01279-9#ref-CR3)</sup>.\n\nPerceptible or not, artistic fingerprints have several useful applications. First, they have practical uses, for example, suggesting possible authors for unattributed works<sup>[1](/articles/s42256-026-01279-9#ref-CR1),[4](/articles/s42256-026-01279-9#ref-CR4),[5](/articles/s42256-026-01279-9#ref-CR5)</sup> or distinguishing authentic pieces from forgeries<sup>[2](/articles/s42256-026-01279-9#ref-CR2)</sup>. Second, they can highlight the networks of influence between artists<sup>[6](/articles/s42256-026-01279-9#ref-CR6),[7](/articles/s42256-026-01279-9#ref-CR7)</sup>, as well as reveal insights into the creative process itself<sup>[8](/articles/s42256-026-01279-9#ref-CR8),[9](/articles/s42256-026-01279-9#ref-CR9)</sup>. For example, the chemical make-up of the paint in an artwork can point to where particular materials were purchased, situating the work within a specific time, place and artistic community. Finally, they can also be useful in educational contexts, helping to train artists to understand the distinctive voices of important figures within their particular discipline.\n\nIdentifying these components has been a central task within many scholarly accounts of visual art<sup>[10](/articles/s42256-026-01279-9#ref-CR10),[11](/articles/s42256-026-01279-9#ref-CR11)</sup>, music<sup>[12](#ref-CR12),[13](#ref-CR13),[14](/articles/s42256-026-01279-9#ref-CR14)</sup> and writing<sup>[15](/articles/s42256-026-01279-9#ref-CR15),[16](/articles/s42256-026-01279-9#ref-CR16)</sup> from the past century and earlier. However, traditional manual methods of analysis are necessarily slow and, hence, hard to apply at scale. Consequently, such research is mostly restricted to a small minority of famous artists from the Western canon, such as Beethoven<sup>[12](/articles/s42256-026-01279-9#ref-CR12)</sup>, Cervantes<sup>[16](/articles/s42256-026-01279-9#ref-CR16)</sup>, Dante<sup>[15](/articles/s42256-026-01279-9#ref-CR15)</sup> and Raphael<sup>[11](/articles/s42256-026-01279-9#ref-CR11)</sup>.\n\nMachine learning provides a more scalable method for learning artistic fingerprints. The standard approach is to train a supervised learning model to identify the creator of an artwork from a given input. Recent deep neural networks perform very well at this task, reaching a high degree of accuracy when identifying visual artists<sup>[17](/articles/s42256-026-01279-9#ref-CR17)</sup>, writers<sup>[4](/articles/s42256-026-01279-9#ref-CR4),[5](/articles/s42256-026-01279-9#ref-CR5),[18](/articles/s42256-026-01279-9#ref-CR18)</sup>, composers<sup>[19](#ref-CR19),[20](#ref-CR20),[21](#ref-CR21),[22](#ref-CR22),[23](/articles/s42256-026-01279-9#ref-CR23)</sup> and performers<sup>[24](#ref-CR24),[25](#ref-CR25),[26](/articles/s42256-026-01279-9#ref-CR26)</sup> from their works. However, such models can be hard for humans to interpret, constraining their general utility for elucidating the creative processes or for contributing to pedagogy. An important current challenge is, therefore, to develop approaches that achieve both high predictive accuracy and are straightforward to interpret.\n\nAlthough this question has been studied extensively in both visual art<sup>[9](/articles/s42256-026-01279-9#ref-CR9),[17](/articles/s42256-026-01279-9#ref-CR17),[27](/articles/s42256-026-01279-9#ref-CR27)</sup> and writing<sup>[1](/articles/s42256-026-01279-9#ref-CR1),[4](/articles/s42256-026-01279-9#ref-CR4),[18](/articles/s42256-026-01279-9#ref-CR18)</sup>, relatively little has been done in the musical domain. Here we address jazz improvisation, which provides a particularly rich manifestation of artistic style, combining three essential aspects of musical creativity: composition (the aspects of the piece that are written in advance), improvisation (how the composed ideas are spontaneously fleshed out and extended) and performance (how these ideas are realized physically, in conjunction with the other musicians). Each of these three elements can become part of a performer’s fingerprint, as has been elucidated in numerous theoretical writings<sup>[14](/articles/s42256-026-01279-9#ref-CR14),[28](#ref-CR28),[29](#ref-CR29),[30](#ref-CR30),[31](/articles/s42256-026-01279-9#ref-CR31)</sup>.\n\nComputational analysis of jazz has historically been difficult. Most jazz performances only exist as audio recordings, and models trained directly on audio are typically hard for humans to interpret<sup>[26](/articles/s42256-026-01279-9#ref-CR26),[32](/articles/s42256-026-01279-9#ref-CR32)</sup>. They may also focus overly on contextual elements (for example, recording quality) rather than the features (harmonic or melodic vocabulary) generally of interest to musicians<sup>[33](/articles/s42256-026-01279-9#ref-CR33),[34](/articles/s42256-026-01279-9#ref-CR34)</sup>. However, recent developments in audio signal processing allow us to bring this work into the symbolic domain<sup>[35](/articles/s42256-026-01279-9#ref-CR35),[36](/articles/s42256-026-01279-9#ref-CR36)</sup>, which supports a much more interpretable analysis.\n\nWe take advantage of these possibilities in the present work, developing models to address many issues relevant to music theorists, performers and educators. These include\n\n- \nWhich musical cues (for example, individual melodic patterns or harmonic progressions) make up a given performer’s fingerprint?\n- \nHow do these cues vary between performance contexts (for example, playing in an ensemble versus unaccompanied)?\n- \nHow are the fingerprints of different performers related?\n- \nHow do machine-extracted authorial cues relate to cues identified by human experts, as published in the existing analytical literature?\n\nTo gain a balanced perspective on these questions and to avoid over-generalizing from one architecture<sup>[37](/articles/s42256-026-01279-9#ref-CR37),[38](/articles/s42256-026-01279-9#ref-CR38)</sup>, we consider a variety of machine learning models in turn. Each model is trained to predict performer identity from a different input representation, across a large dataset of transcribed recordings by 20 famous jazz pianists (see the ‘Dataset’ section). One approach (see the ‘Bag of features approach’ and ‘Training the bag of features approach’ sections) uses large ‘bags’ of basic musical elements (melodic *n*-grams and musical chords); a second approach (see the ‘Representation learning approach’ and ‘Training the representation learning approach’ sections) learns a black-box representation from the music transcription; a third, intermediate approach (see the ‘Multiple input representations approach’ and ‘Training the multiple input representations approach’ sections) learns separate representations for four fundamental musical dimensions, namely, melody, harmony, rhythm and dynamics.\n\nTo contextualize what our models have learned, we interpret their behaviour in light of existing pedagogical and musicological writings. The approach, therefore, provides a way to test claims about authorial cues made in prior literature. Moreover, by validating our models against existing analyses of well-studied artists, we gain confidence that the models could deliver worthwhile analyses when generalized to less well-studied artists. To support such applications in the future, we release open-source implementations and checkpoints for our models, in the hope that they can be used to explore both new and under-researched musical works (see the ‘Code availability’ section).\n\n## Results\n\n### Bag of features approach\n\nWe train three classical supervised learning architectures (logistic regression (LR), random forest (RF) and support vector machine (SVM)) to identify the pianist playing in each recording, using a bag of distinct features, corresponding to individual melodic patterns (*n*-grams) and chord voicings (see the ‘Training the bag of features approach’ section).\n\n#### Bag of features: who is the performer of an unknown musical piece?\n\nWe show the accuracy of the three models in predicting the held-out test data in Extended Data Table [1](/articles/s42256-026-01279-9#Tab1). The LR model (with L2 or ‘ridge’ regularization as the penalty term) performs the best, achieving a top-1 accuracy of 0.767 when predicting the identity of the 20 jazz pianists considered here. The top-5 accuracy for this model is 0.939—meaning that for nearly 95% of unseen recordings, the actual pianist is within the five classes with the highest predictive probability estimated by this model, despite it only using melodic *n*-grams and chord voicings as input features. The RF and SVM models both perform worse, achieving top-1 accuracies of 0.454 and 0.687 on the held-out test data, respectively<sup>[39](/articles/s42256-026-01279-9#ref-CR39)</sup>. Additionally, we find that models trained on chromatic-scale steps perform substantially better than those trained on diatonic-scale steps, even when controlling for the use of mode by each pianist (Supplementary Section [1.1.3](/articles/s42256-026-01279-9#MOESM1)).\n\n#### Bag of features: which domains matter in a musical fingerprint?\n\nOne way of ascertaining the relative importance of the input features used by these models is to calculate the decrease in predictive accuracy when the values obtained for a feature (or group of features) are permuted<sup>[40](/articles/s42256-026-01279-9#ref-CR40)</sup>. We compute feature importance by permuting either all harmony or all melody features and measuring the loss in the held-out test accuracy for the LR versus predicting with the complete feature set. The accuracy loss (averaged over *N* = 1,000 iterations) when permuting melodic features is substantially greater than permuting chord features (0.403 versus 0.152). Similar trends are observed for both optimized RF and SVM models (Extended Data Fig. [1](/articles/s42256-026-01279-9#Fig7)). This might suggest that melodic patterns are more important than chord voicings in differentiating one jazz pianist from another.\n\nGiven the imbalance in the number of features (with over twice as many *n*-grams than voicings), an alternative way to consider the importance of a single feature is to permute random subsets of *K* melodic patterns or chord voicings (sampled without replacement from the full feature space) and measure the mean loss in accuracy. With *K* = 2,000 features, the loss in accuracy for the LR is larger when permuting chord voicings than when permuting melodic patterns (0.015 versus 0.057: mean across *N* = 1,000 iterations). This would suggest that the typical chord voicing has over three times the explanatory power of the typical melody *n*-gram; however, as the vocabulary of melody *n*-grams is so much larger than the vocabulary of voicings, melody *n*-grams still end up more predictive on aggregate than chord voicings<sup>[39](/articles/s42256-026-01279-9#ref-CR39)</sup>.\n\n#### Bag of features: how do musical fingerprints differ between performance contexts?\n\nAnother question we can ask is whether there are differences in how musical features are used in either solo or ensemble performances. To do so, we refit the LR model separately on solo (${{\\mathcal{D}}}^{{\\rm{solo}}}$) and ensemble (${{\\mathcal{D}}}^{{\\rm{trio}}}$) recordings, using the top-*K* melodic and harmonic features, and compute Pearson correlation coefficients between the rows of the resulting weight matrices for every performer (see the ‘Extracting maximally predictive features’ section). In nearly every case, these correlations are positive: for melody, mean(*r*) = 0.267, s.d. = 0.108; for harmony, mean(*r*) = 0.150, s.d. = 0.087 (Fig. [1](/articles/s42256-026-01279-9#Fig1)).\n\nTo the extent that melodic or harmonic cues differ between solo and ensemble performances, we would expect correlations to be lower across the dataset partitions (that is, between ${{\\mathcal{D}}}^{{\\rm{solo}}}$ and ${{\\mathcal{D}}}^{{\\rm{trio}}}$) than within partitions. To test this hypothesis, we conduct a one-sided (left-tailed) permutation-based significance test, which involves randomly shuffling the labels within the dataset before obtaining ${{\\mathcal{D}}}^{{\\rm{solo}}}$ and ${{\\mathcal{D}}}^{{\\rm{trio}}}$. This yields a Monte Carlo null distribution of correlation coefficients for every performer (*N* = 1,000 iterations) against which we test our observed correlations. We apply Bonferroni correction for the number of feature sets (that is, *n* = 2), thereby controlling the false discovery rate on the performer level. These tests indicate that the correlation between feature weights obtained from either dataset are substantially smaller than would be expected if the two datasets were equivalent.\n\n#### Bag of features: which features are associated with particular performers?\n\nAlthough this analysis is interesting in terms of relatively high-level musical musical cues, it does not tell us which features influenced the model to classify individual performers—in other words, which melodic patterns or chord voicings best define the vocabulary of one particular performer. Ref. <sup>[41](/articles/s42256-026-01279-9#ref-CR41)</sup> explored this task by ranking how frequently different jazz improvisers used particular melodic patterns. Similarly, ref. <sup>[42](/articles/s42256-026-01279-9#ref-CR42)</sup> defined the distinctiveness of a pattern within a musical corpus as the degree to which it is over-represented with respect to an anti-corpus.\n\nWe take a different approach by instead formalizing the ‘distinctiveness’ of a given musical feature for a particular performer using the weights it is associated with by our model. This helps mitigate the influence of patterns or voicings that are commonly used by all performers in the dataset, avoiding the need to explicitly define an anti-corpus.\n\nAfter fitting the LR model to the entire dataset, we obtain the matrix **W** (see the ‘Extracting maximally predictive features’ section). Each entry **W**<sub>*i*,*j*</sub> represents the weight assigned to feature *j* ∈ *J* for performer *i* ∈ *I*. A positive weight **W**<sub>*i*,*j*</sub> > 0 indicates that the presence of feature *j* increases the likelihood of predicting performer *i*; conversely, a negative weight **W**<sub>*i*,*j*</sub> < 0 decreases that likelihood. In Fig. [2](/articles/s42256-026-01279-9#Fig2), we show weights for the top- and bottom-five melodic patterns associated with classifications of Bill Evans and Oscar Peterson—the two pianists with the greatest number of individual recordings in the dataset (see the ‘Dataset’ section). We provide similar graphs for all other pianists in Supplementary Figs. [5](/articles/s42256-026-01279-9#MOESM1)–[22](/articles/s42256-026-01279-9#MOESM1), several of whom we discuss in Supplementary Section [1.2](/articles/s42256-026-01279-9#MOESM1). The five most distinctive melodic patterns for each pianist can also be listened to within the context of their performances as part of an interactive web application we have developed (see the ‘Code availability’ section).\n\nFor Bill Evans, two of the five patterns with the strongest weighting outline either a descending major (0, −4, −7, −11) or minor (0, −3, −7, −10) seventh arpeggio, beginning on the seventh and falling to the root (Supplementary Figs. [43](/articles/s42256-026-01279-9#MOESM1) and [44](/articles/s42256-026-01279-9#MOESM1)). Both patterns appear frequently in jazz pianist Jacky Naylor’s pedagogical textbook outlining Evans’ improvisation style<sup>[43](/articles/s42256-026-01279-9#ref-CR43)</sup>. Evans uses this pattern considerably more than any other pianist in our dataset, with it appearing 128 times across his recordings; other pianists who frequently use this pattern include Keith Jarrett (34 appearances) and Brad Mehldau (33).\n\nSeveral of the patterns with the strongest weighting for Oscar Peterson can be considered melodic ‘enclosures’. For example, the pattern (0, −2, −4, −3) is described in ref. <sup>[44](/articles/s42256-026-01279-9#ref-CR44)</sup> on Peterson’s style as a ‘chromatic enclosure around the third’, approaching the major third of the underlying harmony from above and then below (Supplementary Fig. [45](/articles/s42256-026-01279-9#MOESM1)). Again, Peterson uses this pattern the most of any pianist in our dataset (146 appearances). Others who use this pattern frequently include Kenny Barron (94) and Keith Jarrett (97). The (0, 7, 4, 5) pattern associated with Peterson also acts as a ‘scale note enclosure’<sup>[44](/articles/s42256-026-01279-9#ref-CR44)</sup>, approaching the note a perfect fourth above the starting pitch diatonically from above and then below (Supplementary Fig. [46](/articles/s42256-026-01279-9#MOESM1)).\n\nMany of the melodic patterns with strong positive weights are, perhaps, well known to musicians. However, we also wish to draw attention to the patterns with negative weightings: in other words, the features that particular performers do not seem to use. For Evans, these include octave tremolo figures that appear prominently in the ‘bluesy’ playing of both Peterson and Junior Mance (Supplementary Fig. [14](/articles/s42256-026-01279-9#MOESM1)); for Peterson, these patterns include unidirectional scalar patterns that avoid the kinds of disjunct motion typically found in melodic enclosures. In contrast to the features with positive weights—which, as we have already demonstrated, appear often in the pedagogical literature—these negatively weighted patterns could capture avoidance behaviours, which are rarely made explicit in previous work.\n\n#### Bag of features: how are the fingerprints of particular performers related?\n\nThese preceding analyses focused on individual *n*-grams, but it is also possible to use dimensionality reduction to visualize larger-scale clusters of *n*-gram usage. Here we use principal component analysis (PCA; Supplementary Section [1.3](/articles/s42256-026-01279-9#MOESM1)); we apply it to melodic *n*-grams of length 4, of which there are 5,838. In Fig. [3](/articles/s42256-026-01279-9#Fig3), we plot all performer classes and a subset of the melodic patterns that load most strongly onto the first four principal components. In Supplementary Fig. [54](/articles/s42256-026-01279-9#MOESM1), we show the number of times each performer uses the three patterns that load most strongly onto both components 1 and 2. Additional analysis of the remaining components is contained in Supplementary Section [1.3](/articles/s42256-026-01279-9#MOESM1).\n\nMelodic patterns that load positively onto principal component 1 typically consist of small scalar fragments with a span of less than a perfect fifth, such as (0, 1, 3, 5) and (0, 2, 3, 5). Keith Jarrett, Tommy Flanagan and Kenny Barron load positively onto this component. Supplementary Fig. [22](/articles/s42256-026-01279-9#MOESM1) shows how Flanagan uses several of these melodic patterns. Patterns that load negatively onto component 1 span a much larger distance, such as (0, 1, 6, 1) and (0, 1, 2, −9), with Abdullah Ibrahim, Ahmad Jamal and Thelonious Monk also loading negatively here.\n\nMelodic patterns loading positively onto component 2 involve chromatic fragments such as (0, −1, −2, −3) and (0, −1, −3, −2). Bill Evans loads positively onto this component; Supplementary Fig. [43](/articles/s42256-026-01279-9#MOESM1) shows him using (0, −1, −2, −3) to connect sequential appearances of (0, −4, −7, −11). Patterns loading negatively onto component 2 consist of the same interval alternating, including (0, 5, 0, 5), with Brad Mehldau, Junior Mance and McCoy Tyner loading negatively. Many of these negatively loading melodic patterns resemble ostinato figures typically played by a bassist, which may explain Mehldau’s association with them, given that most of his recordings in our dataset are unaccompanied. Most patterns also alternate between perfect intervals, including fourths (0, 5, 0, 5) and octaves (0, 12, 7, 12), associated with Tyner (Supplementary Fig. [18](/articles/s42256-026-01279-9#MOESM1)) and Mance (Supplementary Fig. [14](/articles/s42256-026-01279-9#MOESM1)), respectively.\n\n### Representation learning approach\n\nOne weakness of the bag of features approach is that it is difficult to be certain that a feature ‘bag’ includes all features that might be beneficial for a model. We did not consider, for example, how the arrangement of different chords in sequence might impact voice leading, which is a skill that many jazz pianists devote substantial time to mastering<sup>[29](/articles/s42256-026-01279-9#ref-CR29),[45](/articles/s42256-026-01279-9#ref-CR45),[46](/articles/s42256-026-01279-9#ref-CR46)</sup>.\n\nConsequently, rather than defining the feature space in advance, it may be attractive to allow the model to learn a representation directly from an input, which can then be used in classification. We evaluate two neural network architectures (a convolutional recurrent neural network (CRNN) and a ‘ResNet’) on our dataset. These models are trained directly on the ‘piano rolls’ transcribed from each recording (see the ‘Training the representation learning approach’ section).\n\n#### Representation learning: who is the performer of an unknown musical piece?\n\nWe show the accuracy of each model when predicting the performer of an unseen test recording (Extended Data Table [2](/articles/s42256-026-01279-9#Tab2)). The best-performing model from these experiments is the ResNet trained using our data augmentation pipeline (accuracy, 0.944). This model considerably outperforms both the previous-best LR model (0.767) and the CRNN with augmentation (0.825), and can be considered to set a strong baseline for automatic performer identification models on this dataset.\n\n#### Representation learning: which features are associated with particular performers?\n\nWe use the locally interpretable model explanations (LIME) technique<sup>[47](/articles/s42256-026-01279-9#ref-CR47)</sup> to interpret test-split predictions made by the ResNet trained with augmentation. We use the default settings in the Python library (v. 0.2.0.1) provided by the authors of the original paper. Figure [4](/articles/s42256-026-01279-9#Fig4) highlights the four regions with the strongest LIME attributions (that is, that push the model to positively identify the target class) for clips by Bill Evans and Oscar Peterson. Similar figures for clips by Abdullah Ibrahim, Chick Corea and Keith Jarrett are provided in Supplementary Fig. [55](/articles/s42256-026-01279-9#MOESM1).\n\nThe LIME technique does seem to emphasize musical gestures that could reasonably be considered distinctive for each performer. These include particular ascending and descending melodic patterns for Evans and several chord voicings for Peterson: interestingly, in the case of Evans, one of the attributed regions even contains the descending (0, −4, −7, −11) arpeggio observed previously in Fig. [2a](/articles/s42256-026-01279-9#Fig2). Nonetheless, it is difficult to understand exactly what aspects of those gestures the model is paying attention to. Considering several of the scalar patterns highlighted for Evans, their explanatory power could reasonably derive from the intervallic structure of the scale, the harmony that it outlines, or the rhythmic and dynamic trends that shape it.\n\n### Multiple input representations approach\n\nAs one possible solution to the opaqueness of techniques like LIME, we now evaluate an architecture that learns separate representations for four fundamental musical domains—melody, harmony, rhythm and dynamics<sup>[48](/articles/s42256-026-01279-9#ref-CR48)</sup>—but then combines information from these four domains to make its predictions (see the ‘Training the multiple input representations approach’ section). Each domain is represented as a piano roll with the same shape as the original transcription. We then train small convolutional sub-networks to learn from each representation and aggregate the outputs together to generate a single feature vector that can be used to make predictions (Extended Data Fig. [2](/articles/s42256-026-01279-9#Fig8)), as in ref. <sup>[49](/articles/s42256-026-01279-9#ref-CR49)</sup>.\n\n#### Multiple input representations: who is the performer of an unknown musical piece?\n\nThe results for this approach are shown in Extended Data Table [3](/articles/s42256-026-01279-9#Tab3). The best-performing multi-input model almost matches the performance of the best model described in the ‘Representation learning approach’ section (accuracy of 0.913 versus 0.944). There is, therefore, some trade-off between interpretability and predictive performance, at least in our case. However, it is impressive that the multi-input model can nearly match the baseline set in the ‘Representation learning: who is the performer of an unknown musical piece?’ section, and vastly outperforms the models introduced in the ‘Bag of features: who is the performer of an unknown musical piece?’ section.\n\n#### Multiple input representations: which musical domains contribute the most towards accurate performer identification?\n\nThere are at least two ways to evaluate the high-level musical domains this model is trained on. The first considers how important each representation is to the full model by computing the loss in accuracy when the output of a single sub-network is masked with zeros (Fig. [5a](/articles/s42256-026-01279-9#Fig5)). Masking either the melody, rhythm or harmony sub-network leads to a small, relatively consistent drop in accuracy, with the greatest loss observed for rhythm (rhythm accuracy loss, −0.069; melody, −0.063; harmony, −0.056). However, masking the output of the ‘dynamics’ sub-network leads to a considerably smaller loss in accuracy (−0.019), which suggests that the dynamics are relatively unimportant when predicting the performer of an unknown musical excerpt.\n\nA second way to consider the contributions of each domain is to compute how well a single representation predicts by itself, when the output of the other three sub-networks is masked (Fig. [5b](/articles/s42256-026-01279-9#Fig5)). As before, the predictive accuracy is the lowest when using only the dynamics sub-network (accuracy, 0.263)—and is only slightly better than a model that simply predicts the majority class (0.175). Both melody and rhythm sub-networks yield similar predictive accuracies (0.575 and 0.619, respectively). The most accurate predictions are obtained using only the harmony sub-network, with nearly three-quarters of unseen recordings being classified correctly (0.744). We show the predictive accuracy obtained for all combinations of sub-networks in Supplementary Fig. [56](/articles/s42256-026-01279-9#MOESM1).\n\nThe importance of melodic content for distinguishing between different improvising jazz musicians has been demonstrated previously in the literature on computational modelling<sup>[50](/articles/s42256-026-01279-9#ref-CR50),[51](/articles/s42256-026-01279-9#ref-CR51)</sup>. The same is also true for rhythm<sup>[3](/articles/s42256-026-01279-9#ref-CR3)</sup>, with ref. <sup>[52](/articles/s42256-026-01279-9#ref-CR52)</sup> noting how ‘expressive features of “time feel” serve to define the stylistic profile of [particular] jazz musicians’. Finally, that harmony should be particularly helpful when predicting an unknown performer is perhaps unsurprising: in his monograph on jazz, Gridley<sup>[53](/articles/s42256-026-01279-9#ref-CR53)</sup> describes how ‘each pianist’s particular approach to… chording [is] a signature for [their] style’.\n\n#### Multiple input representations: which domains best distinguish particular performers?\n\nIn Extended Data Figs. [3](/articles/s42256-026-01279-9#Fig9) and [4](/articles/s42256-026-01279-9#Fig10), we show the per-class accuracy for predictions made using a single sub-network, which allows us to consider the degree to which particular domains best distinguish individual performers. For instance, several musicians are particularly distinctive for their use of rhythm. This includes Kenny Barron and Chick Corea, as all of their recordings in the held-out test set can be correctly classified solely using the rhythm sub-network, with considerably lower accuracy obtained when using only the harmony sub-network (0.800 for Kenny Barron and 0.714 for Chick Corea). A possible explanation could be that Corea is one of the only pianists in our dataset to have extensively recorded Latin American music<sup>[54](/articles/s42256-026-01279-9#ref-CR54)</sup>. The distinctiveness of every domain for each performer can be explored as part of our interactive web application (see the ‘Code availability’ section).\n\n## Discussion\n\nThe purpose of this research was to deconstruct the fingerprints of individual artists—the signature elements that make their work recognizable and distinct—by comparing the outputs of machine learning models trained on particular artworks with expert human analyses found in the scholarly literature. This computational approach to the analysis of art has several intrinsic benefits. Human experts can identify fingerprints within small numbers of individual artworks<sup>[28](/articles/s42256-026-01279-9#ref-CR28)</sup>. Our computational modelling, however, allows much larger datasets to be processed (see Fig. [6](/articles/s42256-026-01279-9#Fig6)), and it offers a level of statistical rigour that is unobtainable with manual approaches. Indeed, we find that many of the features highlighted by the model align with prior observations in the theoretical literature, and this provides a useful validation of the approach (for example, see the patterns shown in Fig. [2](/articles/s42256-026-01279-9#Fig2)). However, other important features are not discussed at all in the existing literature, and our models provide straightforward ways to identify them.\n\nOur work also carries pedagogical implications. Critical writing on art often relies on language (for example, ‘swing’, ‘groove’ and ‘feel’ in jazz) that describes particular artists, yet finding clear examples here can be challenging. We have demonstrated that computational models can automatically associate performers with distinctive features, with interesting implications for demystifying their playing. To this end, we have also created a web application (see the ‘Code availability’ section) that enables users to explore the playing of the 20 pianists considered here.\n\nWe can foresee at least two limitations with this work. First, musical dimensions overlap somewhat, making complete separation difficult, as presumed by our multi-input model. For example, harmony is seemingly the most helpful domain for this model when used on its own, whereas rhythm is the most important when all other domains are included (Fig. [5](/articles/s42256-026-01279-9#Fig5)). One possible interpretation could be that some melodic information bled into the harmony representation, and vice versa. Although perfect separation may be impossible, treating these as distinct constructs can yield insights harder to obtain from fully entangled representations, as demonstrated in the LIME analysis (Fig. [4](/articles/s42256-026-01279-9#Fig4)).\n\nSecond, the inputs themselves carry limitations. MIDI piano rolls are widely used by musicians and lend themselves to meaningful interpretation—unlike auditory representations such as spectrograms<sup>[32](/articles/s42256-026-01279-9#ref-CR32)</sup>. However, they necessarily constrain the musical information that can be learned. Although this can mitigate any potential overfitting to recording- or performer-level characteristics (Supplementary Section [1.4](/articles/s42256-026-01279-9#MOESM1)), jazz improvisation engages with tonal and timbral qualities that MIDI cannot easily capture, like vibrato and pitch-bending<sup>[51](/articles/s42256-026-01279-9#ref-CR51)</sup>. A revised pipeline applied to instruments such as the saxophone would need to engage with these musical techniques. Additionally, extending the pipeline to double bass<sup>[55](/articles/s42256-026-01279-9#ref-CR55)</sup>, for instance, would allow us to examine how that instrument shapes the harmonic function of the chord voicings (see the ‘Bag of features approach’ section).\n\nThis work nonetheless unlocks several exciting avenues for further research. Beyond individual artists, fingerprints can be examined through geographical, historical and cultural lenses—for instance, the ‘Detroit school’ of pianists, including Tommy Flanagan and Hank Jones<sup>[54](/articles/s42256-026-01279-9#ref-CR54)</sup>. Our models could investigate such networks by studying the evolution and transmission of distinctive musical cues over time or across locations<sup>[7](/articles/s42256-026-01279-9#ref-CR7),[56](/articles/s42256-026-01279-9#ref-CR56),[57](/articles/s42256-026-01279-9#ref-CR57)</sup>. Training equivalent models on classical performances would enable cross-genre comparisons<sup>[58](/articles/s42256-026-01279-9#ref-CR58)</sup>. Alternative methods—such as concept-based techniques<sup>[32](/articles/s42256-026-01279-9#ref-CR32)</sup>—could also be applied to our multi-input model’s representations. Future research could also be conducted into improving methods for automatic scale estimation in polyphonic jazz (Supplementary Section [1.1](/articles/s42256-026-01279-9#MOESM1)), which could improve the performance of models trained on diatonic (versus chromatic) scale steps (Extended Data Table [1](/articles/s42256-026-01279-9#Tab1)). Finally, our models could be extended to less well known or historically under-represented musicians.\n\n## Methods\n\n### Dataset\n\nOur dataset consists of the recordings of jazz piano improvisation by 20 famous performers, transcribed using an automatic system into MIDI ‘piano roll’ format. We study both group and solo performances, which are known to differ in systematic ways. For instance, in jazz, the piano shares responsibility with the bass and drums for defining the harmonic and rhythmic movement of a performance<sup>[28](/articles/s42256-026-01279-9#ref-CR28)</sup>. When a pianist performs unaccompanied, they may need to compensate for the absence of the other instruments, such as by emphasizing harmonically ‘fuller’ chords or bass lines<sup>[29](/articles/s42256-026-01279-9#ref-CR29),[45](/articles/s42256-026-01279-9#ref-CR45)</sup>.\n\n#### Source datasets\n\nWe use transcriptions from two existing open-source datasets: (1) the jazz trio database (JTD)<sup>[59](/articles/s42256-026-01279-9#ref-CR59)</sup> and (2) the piano jazz with automatic MIDI annotations<sup>[33](/articles/s42256-026-01279-9#ref-CR33)</sup> dataset. JTD contains transcriptions of 1,294 improvised solos by 34 different jazz pianists performing in a trio with a bassist and drummer. Suitable performers were identified from online listening data and pedagogical textbooks, with all recordings from their discographies that met a predefined inclusion criteria (for example, instrumentation and musical attributes) included in the dataset. PiJAMA contains transcriptions of 2,777 full-length performances by 120 different pianists, without accompaniment. Suitable performers were identified from both textbooks and records of finalists in international jazz competitions, with all available performances by these pianists included in the dataset. Note that as JTD and PiJAMA contain recordings of different types of jazz performance (solo and trio) identified manually by the dataset creators, combining both together is highly unlikely to introduce contamination (that is, same recordings contained in both datasets).\n\n#### Transcription and preprocessing\n\nThe transcriptions for both datasets were created by applying the automatic music transcription system described by ref. <sup>[36](/articles/s42256-026-01279-9#ref-CR36)</sup> to an audio signal. This model was the state-of-the-art system made available openly at the time both datasets were compiled. For the recordings in PiJAMA, an automatic music tagging system was first used to filter out non-music portions of the audio (for example, applause and spoken introductions), with the transcription model applied to the remaining sections. For the recordings in JTD, the audio was manually trimmed to the piano solo; a state-of-the-art, open-source instrument separation model<sup>[60](/articles/s42256-026-01279-9#ref-CR60)</sup> was applied to isolate the playing of the pianist; and the transcription model was then applied to the separated audio. The transcriptions in both datasets were created at a resolution of 100 frames per second (that is, 10 ms per frame), the default setting for the transcription model.\n\nReference <sup>[33](/articles/s42256-026-01279-9#ref-CR33)</sup> found that the transcription pipeline used in PiJAMA obtained an average note *F*<sub>1</sub> score of 0.87, compared with a small corpus of hand-curated jazz piano transcriptions. Reference <sup>[39](/articles/s42256-026-01279-9#ref-CR39)</sup> found an average note onset *F*<sub>1</sub> score for JTD of 0.82, compared with a small hand-annotated sample of recordings from this dataset. Both results indicate good performance for an automatic piano transcription model and are similar to those obtained when applying the same pipeline to datasets of Western classical music<sup>[33](/articles/s42256-026-01279-9#ref-CR33)</sup>.\n\nAdditionally, to counter several of the common failure states of automatic transcription models, we preprocess our dataset. We remove any notes with a pitch outside the standard range of the piano (that is, lower than 21 and higher than 108) or that have a duration shorter than 10 ms, which were generally artefacts. We merge together notes played at the same pitch, but with less than 10 ms between successive offset and onset times (‘false triggers’). Finally, we clamp notes to a maximum duration of 5 s—which proved necessary due to a minority of cases (0.22% of notes in the dataset) in which the automatic transcription model failed to correctly apply note-off events. Further preprocessing and heuristics specific to each model are described in the ‘Bag of features approach’, ‘Representation learning approach’ and ‘Multiple input representations approach’ sections.\n\nFinally, we note that the annotations in JTD and PiJAMA are solely restricted to performance MIDI. By contrast, the Weimar jazz database<sup>[61](/articles/s42256-026-01279-9#ref-CR61)</sup> contains additional annotations, including chord symbols and section annotations. Our datasets make up for this, however, with their substantially larger size—with PiJAMA alone being over six times the size of WJD.\n\n#### Data splits\n\nTwenty-five different pianists have at least one recording in both datasets (Fig. [6a](/articles/s42256-026-01279-9#Fig6)). However, the cumulative duration of all recordings varies substantially between performers, from less than half an hour to over nine hours (Fig. [6b](/articles/s42256-026-01279-9#Fig6)). Consequently, we remove pianists with less than 80 min of recordings across JTD and PiJAMA. This leaves 1,629 performances by 20 pianists, with a total duration of 84 h. Our dataset includes a greater number of recordings from JTD (1,001) than PiJAMA (628). However, the PiJAMA recordings (which are all full-length performances) are typically longer than those in JTD (which consist only of the piano solo in a recording), with a median duration of 4 min 18 s versus 1 min 51 s (Fig. [6c](/articles/s42256-026-01279-9#Fig6)).\n\nWe randomly split these recordings into training, validation and test subsets in the ratio of 8:1:1. These splits are stratified by source database, such that a proportional number of solo and ensemble performances are included in each subset. We use these splits to train and evaluate all of the models described in the remainder of the paper. Note that we find no evidence for either the ‘album’ or ‘composition’ effect<sup>[34](/articles/s42256-026-01279-9#ref-CR34)</sup> when training our models on these splits (Supplementary Section [1.4](/articles/s42256-026-01279-9#MOESM1)).\n\nFor the models described in the ‘Bag of features approach’ section, features are extracted from an entire recording. For the neural networks in the ‘Representation learning approach’ and ‘Multiple input representations approach’ sections, as in ref. <sup>[19](/articles/s42256-026-01279-9#ref-CR19)</sup>, we first segment each recording into 30-s clips (see the ‘Dataset’ section), with each clip inheriting the performer label of its parent recording during training. The hop size for each clip is 30 s during inference and varies between 15 and 30 s during training, as part of our data augmentation pipeline (see the ‘Data augmentation’ section). During training, the classification loss is calculated using class probabilities estimated separately from each clip. During evaluation, we produce a track-level accuracy metric comparable with the models trained on entire recordings in the ‘Bag of features approach’ section by averaging the class probabilities estimated across all clips taken from a single parent recording<sup>[19](/articles/s42256-026-01279-9#ref-CR19)</sup>.\n\n### Training the bag of features approach\n\nIn the ‘Bag of features approach’ section, we train three classical supervised learning architectures to identify the pianist playing in each recording using a bag of distinct, identifiable features. These models are RF, SVMs with a linear kernel, and (regularized) multinomial LR. All three architectures have been widely used in previous computational models of musical performances<sup>[3](/articles/s42256-026-01279-9#ref-CR3),[26](/articles/s42256-026-01279-9#ref-CR26),[62](#ref-CR62),[63](#ref-CR63),[64](#ref-CR64),[65](#ref-CR65),[66](/articles/s42256-026-01279-9#ref-CR66)</sup>. The implementations are taken from the scikit-learn (v. 1.5.1) Python library<sup>[67](/articles/s42256-026-01279-9#ref-CR67)</sup>. Note that every model is fit using the ‘one-versus-all’ strategy.\n\n#### Feature extraction\n\nAlthough these models could be trained on a wide variety of predictive features, we restrict ourselves to melodic patterns and chord voicings, as these can be extracted simply from a musical transcription<sup>[68](/articles/s42256-026-01279-9#ref-CR68)</sup>. For an analogous approach studying rhythmic features obtained from many of the same recordings considered here, see ref. <sup>[3](/articles/s42256-026-01279-9#ref-CR3)</sup>.\n\nExtracting melody from a symbolic representation of a polyphonic musical performance is a challenging task. Ideally, we would use a sophisticated algorithm that either simulates underlying cognitive procedures involved in melody perception<sup>[69](/articles/s42256-026-01279-9#ref-CR69)</sup> or that learns to identify melodies from large annotated corpora of polyphonic music. Some deep learning architectures do exist that attempt the latter task<sup>[70](/articles/s42256-026-01279-9#ref-CR70)</sup>, but we find that they perform badly on our data, with very few melody notes identified successfully and the majority of the output being empty (Extended Data Fig. [5](/articles/s42256-026-01279-9#Fig11) and Supplementary Fig. [1](/articles/s42256-026-01279-9#MOESM1)). A possible explanation is that the data used to train these models often does not include jazz piano performances.\n\nA simpler method is the ‘skyline’ algorithm outlined in ref. <sup>[71](/articles/s42256-026-01279-9#ref-CR71)</sup>, which extracts the note with the highest pitch among the concurrent notes played at every unique onset time. A similar method was previously used in ref. <sup>[72](/articles/s42256-026-01279-9#ref-CR72)</sup> to extract melodic content from polyphonic jazz piano performances. The skyline algorithm has drawbacks—for instance, the extracted notes may ‘leap’ between accompaniment and melody. Nevertheless, when comparing the results of the algorithm with a random set of clips annotated manually by a jazz expert, we found strong agreement (Supplementary Section [1.5](/articles/s42256-026-01279-9#MOESM1)).\n\nBefore applying the algorithm, we quantize a transcription by ‘snapping’ the note onsets to the nearest 100 ms (that is, ten frames). This value is roughly equivalent to the duration of a triplet eighth note at the mean tempo of the recordings contained in the JTD<sup>[59](/articles/s42256-026-01279-9#ref-CR59)</sup>. We then apply the skyline and obtain a vector of pitch classes, from which we extract *n*-grams—that is, melodic ‘chunks’ obtained over a sliding window. In this way, we ensure that the length of the extracted *n*-grams does not directly correlate with the tempo of a performance (which would be the case if, for instance, the duration of the window was set to a fixed number of seconds). A similar approach to ours was previously used in ref. <sup>[73](/articles/s42256-026-01279-9#ref-CR73)</sup>.\n\nWe then use heuristics to filter the extracted melodies, mitigating both the drawbacks of the skyline algorithm and several of the common failure states of automatic transcription models (see the ‘Dataset’ section). We delete *n*-grams that have a total span of more than 12 semitones. This removes cases where the skyline probably ‘leaps’ between melody and accompaniment, or where ‘octave errors’ occur in the output of the transcription model. Additionally, we discount appearances of *n*-grams with at least one duration greater than 2 s between successive offsets and onsets, before quantization. This removes cases where *n*-grams might span across phrase boundaries, or where notes are not detected by the transcription model. Note that 2 s is the approximate duration of one measure at the average tempo of the recordings contained in the JTD<sup>[59](/articles/s42256-026-01279-9#ref-CR59)</sup>. Finally, we convert each *n*-gram into a transposition-invariant representation by computing the difference between successive pitch classes.\n\nWe use values of *n* ∈ {3, 4, 5, 6, 7} when extracting melodic patterns from each transcription. In Extended Data Fig. [6](/articles/s42256-026-01279-9#Fig12), we empirically demonstrate that these values of *n* are sufficient to achieve the ceiling accuracy for the LR model when predicting the held-out validation split of the dataset, and that using larger values of either min(*n*) or max(*n*) does not increase performance<sup>[62](/articles/s42256-026-01279-9#ref-CR62)</sup>. However, this does mean that some of our melodies obtained with lower values of *n* can more accurately be described as short patterns, rather than full-length phrases. Nonetheless, we note that learning such melodic ‘chunks’ does often form a key part of jazz pedagogy<sup>[29](/articles/s42256-026-01279-9#ref-CR29)</sup>. For an analogous procedure that solely considers longer melodic patterns in jazz, see ref. <sup>[41](/articles/s42256-026-01279-9#ref-CR41)</sup>.\n\nTo extract chord voicings from the MIDI transcription, we follow a method similar to ref. <sup>[74](/articles/s42256-026-01279-9#ref-CR74)</sup>. First, we quantize the transcription to the nearest 100 ms according to the onset time of each note, as before. We keep quantized frames that contain *n* ∈ {3, 4, 5, 6, 7} notes from each performance—in other words, chord voicings that contain between three (that is, triads) and seven notes. We discard chords with two or more leaps of greater than 15 semitones between two adjacent pitches in the chord, since such chords are more-or-less unplayable with an ordinary handspan, and so are likely to correspond to transcription errors instead. Finally, we convert each chord into a transposition-invariant representation by expressing every pitch in terms of the number of semitones it lies above the lowest note in the chord. Note that we choose not to subtract the skylined melody as the highest note of every chord could theoretically have both a harmonic and melodic function, such as in the ‘locked hands’ mode of a jazz piano performance<sup>[45](/articles/s42256-026-01279-9#ref-CR45)</sup>.\n\nWe obtain counts for a total of 430,841 melodic patterns and 53,198 chord voicings (total features, 484,039) for the 1,629 recordings in the dataset. In Supplementary Fig. [2](/articles/s42256-026-01279-9#MOESM1), we show the 50 most common 4-grams ranked by their total number of occurrences across the entire dataset. In Supplementary Fig. [3](/articles/s42256-026-01279-9#MOESM1), we show the fifty 4-grams that occur in the greatest number of distinct recordings.\n\nSimilar to how text-based authorship models typically remove both frequent (that is, ‘stop’) and infrequent words<sup>[75](/articles/s42256-026-01279-9#ref-CR75)</sup>, we then discard features contained in fewer than 10 and more than 1,000 recordings to reduce the size of the vocabulary. In Supplementary Fig. [4](/articles/s42256-026-01279-9#MOESM1), we demonstrate that these thresholds appear to be optimal for our dataset, both when compared with alternate values and with no threshold.\n\nThis leaves 17,918 melodic patterns and 3,752 chord voicings (total features, 21,670), with the reduced number of features explainable by a large number of patterns and voicings that appear in very few recordings. We transform the matrix of feature counts to a normalized representation using the term frequency-inverse document frequency method, previously used for composer classification in ref. <sup>[62](/articles/s42256-026-01279-9#ref-CR62)</sup>. We find that the term frequency-inverse document frequency substantially improves the predictive accuracy compared with other techniques such as *z* transformation.\n\nWe show counts for the 20 chord voicings and melodic patterns that occur most frequently after filtering (Extended Data Fig. [7](/articles/s42256-026-01279-9#Fig13)). Each pattern or voicing can be interpreted as showing the number of semitones away from an initial starting pitch, such that the feature (0, −1, −2, −3) becomes pitches (C, B, B*♭*, A) or alternatively (G, G*♭*, F, E). The melodic patterns that occur most often are either alternations of ‘perfect’ intervals (such as the octave (0, −12, 0) and fifth (0, −7, 0)) or scalar fragments (such as a descending chromatic scale (0, −1, −2, −3) or ascending minor scale (0, 2, 3, 5)). As might be expected, shorter patterns appear more frequently than longer ones, with the 20 most common patterns only containing three or four notes. Several of the chord voicings that occur most frequently can be understood with relation to familiar chord types, such as major (0, 4, 7) and minor (0, 3, 7) triads in the root position. Both first (0, 3, 8) and second (0, 5, 9) inversions of the major triad are also common, as are ‘stacked’ octave (0, 12, 24) and perfect fourth (0, 5, 10) intervals.\n\n#### Training\n\nThe process used to train each of the three models is identical to that in ref. <sup>[3](/articles/s42256-026-01279-9#ref-CR3)</sup>. We generate a two-dimensional hyperparameter space (Supplementary Tables [1](/articles/s42256-026-01279-9#MOESM1)–[3](/articles/s42256-026-01279-9#MOESM1)), randomly sample possible configurations of parameters from this space (with *N* = 1,000 iterations), fit the model to the training split of the dataset and measure the accuracy when predicting class labels for the validation split. We find that this number of iterations is sufficient to achieve optimal performance on the held-out validation data across all classifier types. We also find that using a Bayesian optimization technique<sup>[76](/articles/s42256-026-01279-9#ref-CR76)</sup> to tune hyperparameters achieves the same results as random sampling.\n\nWe then use the hyperparameter configuration that resulted in the highest validation accuracy to predict the class labels of the held-out test split, and report the accuracy for this subset of the data as the overall performance of the model. Following ref. <sup>[3](/articles/s42256-026-01279-9#ref-CR3)</sup>—and, for consistency with the neural networks we introduce later—we do not retrain the model on the training and validation splits after selecting the optimal hyperparameter configuration. A single iteration takes approximately 20 s for the LR model, 211 s for the RF model and 31 s for the SVM model. All iterations run in parallel on a machine with 76 CPU cores.\n\n#### Extracting maximally predictive features\n\nIn Section xx, we want to obtain both the most strongly predictive sets of *K* melodic patterns and *K* chord voicings, across all performers. To do so, we first fit the LR model to the entire dataset, using the recordings in the training, validation and testing splits. Note that as we fit this model using the one-versus-all strategy, we have a single weight per feature for every performer. This can be written as the matrix ${\\bf{W}}\\in {{\\mathbb{R}}}^{I\\times J}$, where *I* is the number of performers (20) and *J* is the number of features (21,670).\n\nNext, we compute the absolute magnitude of each feature weight for every performer, and take the maximum across performers. This yields the vector of feature weights ${\\bf{X}}\\in {{\\mathbb{R}}}^{J}$, calculated as ${{\\bf{X}}}_{j}={\\mathop{\\max }\\nolimits_{i\\in I}}\\left\\vert {{\\bf{W}}}_{i,j}\\right\\vert ,\\,\\forall j\\in J$. Next, we obtain the indices for both the top-*K* melodic patterns and top-*K* chord voicings (with *K* = 2,000, as before) by sorting **X**. This lets us obtain two separate subsets of features *j* ∈ *J*, each with length *K*: ${{\\mathcal{F}}}^{{\\rm{mel}}}$ for melodic patterns and ${{\\mathcal{F}}}^{{\\rm{har}}}$ for chord voicings. Note that although it would be possible to use all patterns or voicings here, many of these would likely be either redundant or non-discriminative given the high dimensionality of the feature space, which could artificially deflate the correlations.\n\nSimilarly, we partition the full dataset into two subsets by recording context: ${{\\mathcal{D}}}^{{\\rm{solo}}}$ for recordings in PiJAMA, and ${{\\mathcal{D}}}^{{\\rm{trio}}}$ for recordings in JTD. We then refit the LR model separately on ${{\\mathcal{D}}}^{{\\rm{solo}}}$ and ${{\\mathcal{D}}}^{{\\rm{trio}}}$, once using ${{\\mathcal{F}}}^{{\\rm{mel}}}$ as the feature space and once using ${{\\mathcal{F}}}^{{\\rm{har}}}$, yielding four weight matrices: ${{\\bf{M}}}^{{\\rm{solo}}},{{\\bf{M}}}^{{\\rm{trio}}}\\in {{\\mathbb{R}}}^{I\\times K}$ for melodic patterns and ${{\\bf{H}}}^{{\\rm{solo}}},{{\\bf{H}}}^{{\\rm{trio}}}\\in {{\\mathbb{R}}}^{I\\times K}$ for chord voicings.\n\nFor every performer *i* ∈ *I*, we then compute the Pearson correlation coefficient between the rows of the solo and trio weight matrices, separately for each feature type. Thus, ${r}_{i}^{\\,\\text{mel}}=\\text{corr}\\,\\left({{\\bf{M}}}_{i}^{\\,\\text{solo}\\,},\\ {{\\bf{M}}}_{i}^{\\,\\text{trio}\\,}\\right)$, where ${{\\bf{M}}}_{i}^{\\,\\text{solo}\\,}$ and ${{\\bf{M}}}_{i}^{\\,\\text{trio}\\,}$ denotes the *i*th row of both matrices. A value of ${r}_{i}^{\\,\\text{mel}\\,}$ above 0 indicates that performer *i* uses melodic patterns similarly regardless of whether they are playing solo or with an ensemble; on the contrary, a value close to or below 0 suggests a meaningful divergence in musical cues across the two contexts.\n\n### Training the representation learning approach\n\nIn the ‘Representation learning approach’ section, we evaluate two convolutional neural networks on our dataset. This type of network possesses a range of inductive biases that makes it effective at modelling musical performances in the symbolic domain. For instance, the convolution operation is equivariant to translation; if an input is shifted (for example, if a pitch pattern is transposed), so will the corresponding feature map. When applied to piano roll representations, which put time on the horizontal axis and pitch on the vertical axis, convolutional neural networks, therefore, capture both transposition invariance (the identity of a pitch pattern is preserved when played higher or lower) and temporal invariance (the identity of a pitch pattern is preserved when played earlier or later in a recording). Convolutional neural networks have outperformed both transformer and graph neural network architectures when identifying classical pianists from symbolic representations of their performances<sup>[23](/articles/s42256-026-01279-9#ref-CR23)</sup>.\n\nThe first model we test (CRNN) uses eight convolutional layers and a bidirectional gated recurrent unit followed by a fully connected layer and softmax function to generate class probabilities, as in ref. <sup>[19](/articles/s42256-026-01279-9#ref-CR19)</sup>. The second (ResNet) is the ResNet-50 model, which has previously been used for composer identification in symbolic music<sup>[20](/articles/s42256-026-01279-9#ref-CR20),[32](/articles/s42256-026-01279-9#ref-CR32)</sup>, for identifying ‘samples’ in music audio recordings<sup>[77](/articles/s42256-026-01279-9#ref-CR77)</sup>, and for identifying artists from hand-drawn sketches<sup>[17](/articles/s42256-026-01279-9#ref-CR17)</sup>. Both models are implemented using the PyTorch (v. 2.4.1) Python library<sup>[78](/articles/s42256-026-01279-9#ref-CR78)</sup>. Further description of both models is provided in Supplementary Section [1.6](/articles/s42256-026-01279-9#MOESM1).\n\n#### Input representation\n\nFollowing ref. <sup>[19](/articles/s42256-026-01279-9#ref-CR19)</sup>, we train these models on 30-s clips extracted from every recording. Each clip is represented as a one-channel ‘image’ of shape (88, 3,000), where the channel dimension is the velocity of each note, scaled between 0 and 1 (where a value of 1 is equivalent to the note with the highest velocity in the clip). This has been shown to mitigate overfitting to any variation in dynamic range compression apparent between recordings<sup>[33](/articles/s42256-026-01279-9#ref-CR33),[34](/articles/s42256-026-01279-9#ref-CR34),[79](/articles/s42256-026-01279-9#ref-CR79)</sup>.\n\nThe height of the input corresponds to the pitch range of the piano, with one piano key per bin; the width is the duration of the clip multiplied by the number of frames per second used by the transcription model. We do not downsample the width of the input due to the importance of microtiming in prior analyses of expressive jazz performances<sup>[3](/articles/s42256-026-01279-9#ref-CR3),[26](/articles/s42256-026-01279-9#ref-CR26),[52](/articles/s42256-026-01279-9#ref-CR52),[64](/articles/s42256-026-01279-9#ref-CR64)</sup>.\n\n#### Data augmentation\n\nWe apply data augmentation to the recordings in our training split. Our augmentation pipeline jointly manipulates the pitch, time (that is, onset and offset) and velocity parameters of each note, as well as the hop between two successive 30-s clips (Supplementary Section [1.7](/articles/s42256-026-01279-9#MOESM1)). We assign a 50% probability that data augmentation will be applied to each clip seen during a single training epoch; in this way, our models learn both absolute and relative distinctions between pitch classes. Additionally, data augmentation typically improves the generalizability of a performer identification model<sup>[22](/articles/s42256-026-01279-9#ref-CR22)</sup>, which will be beneficial if our models are to be applied to unseen data—for example, in educational contexts.\n\n#### Training\n\nWe train two versions of both models, with and without data augmentation. Each model is trained for 100 epochs on a single NVIDIA A100 GPU with categorical cross-entropy loss, which we noted was sufficient to minimize the validation loss. With the exception of batch size (which we set to 20), we tune hyperparameters for both models separately. For training the CRNN, we set the learning rate to 0.001 and use the Adam optimizer; for the ResNet, we use the hyperparameter configuration described in ref. <sup>[20](/articles/s42256-026-01279-9#ref-CR20)</sup>. After training, we use the saved model weights with the best accuracy on the validation split to predict the held-out test data, for compatibility with the process described earlier in the ‘Training the bag of features approach’ section. Training takes approximately 5 h for the ResNet and 6.5 h for the CRNN.\n\n### Training the multiple input representations approach\n\nIn the ‘Multiple input representations approach’ section, we evaluate an architecture that learns to identify the pianist in a recording from four separate musical inputs, each representing the melodic, harmonic, rhythmic and dynamic components of a recording<sup>[39](/articles/s42256-026-01279-9#ref-CR39)</sup>.\n\n#### Input representations\n\nGiven a one-channel piano roll with the dimensionality (88, 3,000), we want to split this into four separate rolls, all with the same height and width but which contain the melodic, harmonic, rhythmic or dynamic content from the input. Although it would be possible to use different dimensions for each roll (for instance, representing each melody note using a single pixel), maintaining the same dimensionality ensures that each roll remains directly comparable with the original input and to one another (receiving, for instance, the same number of learnable parameters). We provide a brief description below of how each roll is created; further information is given in Supplementary Section [1.8](/articles/s42256-026-01279-9#MOESM1).\n\nMelody rolls are generated by quantizing and applying the skyline algorithm to the transcription (see the ‘Feature extraction’ section), setting the velocity of each note to binary, and adjusting the inter-onset interval for each note to a consistent value dependent on the total number of melody notes and the width of the input piano roll. Harmony rolls are generated by quantizing the transcription, keeping quantized frames with at least three notes, adjusting inter-onset intervals for notes within one bin, and binarizing the velocity of every note. Rhythm rolls maintain the onset and offset times for each note, but randomize the pitch (that is, height) and binarize the velocity. Dynamics rolls are generated by binning notes by onset time, setting their inter-onset intervals to a consistent value and randomizing their pitch—that is, maintaining only the initial velocity of each note.\n\n#### Model architecture\n\nOur proposed architecture uses four convolutional sub-networks to generate an *x*-dimensional embedding from every input roll. Each sub-network shares the same architecture, but the weights are updated separately. We test a range of architectures in our experiments, including 50-, 34- and 18-layer ResNets—the latter being the smallest implemented in PyTorch. We also test the CRNN architecture described in Supplementary Section [1.6](/articles/s42256-026-01279-9#MOESM1). As a form of regularization, during training, we experiment with masking the outputs of particular sub-networks with zeros. There is a 10% chance that one, two or three sub-networks will be masked when processing a clip (that is, a 30% total probability of masking any number of sub-networks), with 70% probability that all four sub-networks will be used during prediction.\n\nThe outputs from each network are then stacked vertically to create an array with the shape (4, *x*). It would be straightforward to generate the final *x*-dimensional feature vector by either taking the average or maximum of the neurons in every ‘column’ of this array. However, this might not be sufficient to allow the model to capture interactions between different musical domains. Consequently, we experiment with using a self-attention layer (with between 4 and 16 heads) across embeddings before pooling. After aggregation, the resulting *x*-dimensional feature vector is fed into a fully connected layer (with dropout at a rate of 50%) and the softmax function to generate class probabilities. Extended Data Fig. [2](/articles/s42256-026-01279-9#Fig8) shows an outline of the proposed architecture.\n\n#### Training\n\nWe train the model for 100 epochs with a batch size of 20, using 30-s clips and aggregating class probabilities to obtain track-level predictions. We use the Adam optimizer with categorical cross-entropy loss and a learning rate of 0.0001. Training takes between 12 and 21 h on a single NVIDIA A100 GPU, depending on the size of the model. We make a model card<sup>[80](/articles/s42256-026-01279-9#ref-CR80)</sup> available on the same open-source repository that contains the model code (see the ‘Code availability’ section).\n\nWe show the results for our experiments in Extended Data Table [3](/articles/s42256-026-01279-9#Tab3). The most accurate predictions of the held-out test recordings are obtained using the smallest ResNet with 18 layers; increasing the size of each sub-network beyond this decreases the performance, which would suggest that overparameterization occurs with larger architectures<sup>[23](/articles/s42256-026-01279-9#ref-CR23)</sup>. We find that accuracy does not improve with the addition of self-attention. This could imply that there are no meaningful interactions between the different musical domains as they are represented here. As the network has only a single fully connected layer, this would instead imply that the contributions of each sub-network are additive. Finally, we observe slight performance improvements from using masking.\n\n### Reporting summary\n\nFurther information on research design is available in the [Nature Portfolio Reporting Summary](/articles/s42256-026-01279-9#MOESM2) linked to this article.\n\n## Data availability\n\nThe data that support the findings of this study are from two published datasets of symbolic musical transcriptions<sup>[33](/articles/s42256-026-01279-9#ref-CR33),[59](/articles/s42256-026-01279-9#ref-CR59)</sup>. The preprocessed versions of the data and the checkpoints for the trained models are available (with no restrictions on availability) via Zenodo at [https://doi.org/10.5281/zenodo.14774191](https://doi.org/10.5281/zenodo.14774191) (ref. <sup>[81](/articles/s42256-026-01279-9#ref-CR81)</sup>).\n\n## Code availability\n\nThe code for this work is available via Zenodo at [https://doi.org/10.5281/zenodo.20345004](https://doi.org/10.5281/zenodo.20345004) (ref. <sup>[82](/articles/s42256-026-01279-9#ref-CR82)</sup>). The Aeberscale algorithm developed for this work (Supplementary Section [1.1](/articles/s42256-026-01279-9#MOESM1)) is available via Zenodo at [https://doi.org/10.5281/zenodo.20345077](https://doi.org/10.5281/zenodo.20345077) (ref. <sup>[83](/articles/s42256-026-01279-9#ref-CR83)</sup>). The web application developed as part of this work is accessible at [https://cms.mus.cam.ac.uk/jazz-piano-fingerprints-ml](https://cms.mus.cam.ac.uk/jazz-piano-fingerprints-ml) and the relevant code is available via Zenodo at [https://doi.org/10.5281/zenodo.20345004](https://doi.org/10.5281/zenodo.20345004) (ref. <sup>[82](/articles/s42256-026-01279-9#ref-CR82)</sup>).\n\n## References\n\n1. Burrows, J. F. Delta’: a measure of stylistic difference and a guide to likely authorship. *Lit. Linguist. Comput.***17** , 267–287 (2002).\n2. Łydżba-Kopczyńska, B. I. & Szwabiński, J. Attribution markers and data mining in art authentication. *Molecules***27** , 70 (2022).\n3. Cheston, H., Schlichting, J. L., Cross, I. & Harrison, P. M. C. Rhythmic qualities of jazz improvisation predict performer identity and style in source-separated audio recordings. *R. Soc. Open Sci.***11** , 240920 (2024).\n4. Setzu, M., Corbara, S., Monreale, A., Moreo, A. & Sebastiani, F. Explainable authorship identification in cultural heritage applications. *ACM J. Comput. Cult. Herit.***17** , 44 (2024).\n5. Abbasi, A. et al. Authorship identification using ensemble learning. *Sci. Rep.***12** , 9537 (2022).\n6. Abry, P., Wendt, H. & Jaffard, S. When van Gogh meets Mandelbrot: multifractal classification of painting's texture. *Signal Process.***93** , 554–572 (2013).\n7. Frieler, K., Höger, F. & Pfleiderer, M. Anatomy of a lick: structure and variants, history and transmission. In *Book of Abstracts of the Digital Humanities Conference* (2019).\n8. Li, J., Yao, L., Hendriks, E. & Wang, J. Z. Rhythmic brushstrokes distinguish van Gogh from his contemporaries: findings via automated brushstroke extraction. *IEEE Trans. Pattern Anal. Mach. Intell.***34** , 1159–1176 (2012).\n9. Van Noord, N., Hendriks, E. & Postma, E. Toward discovery of the artist's style: learning to recognize artists by their artworks. *IEEE Signal Process. Mag.***32** , 46–54 (2015).\n10. Wölfflin, H. *Principles of Art History: The Problem of the Development of Style in Later Art* (Dover Publications, 1950).\n11. Morelli, G. *Italian Painters: Critical Studies of Their Works* (John Murray, 1892).\n12. Tymoczko, D. *Tonality: An Owner's Manual* (Oxford Univ. Press, 2023).\n13. Hepokoski, J. A. & Darcy, W. *Elements of Sonata Theory: Norms, Types, and Deformations in the Late-Eighteenth-Century Sonata* (Oxford Univ. Press, 2006).\n14. Cugny, L. *Analysis of Jazz: A Comprehensive Approach* (Univ. Press of Mississippi, 2024).\n15. Bloom, H. *The Anxiety of Influence: A Theory of Poetry* (Oxford Univ. Press, 1973).\n16. Spitzer, L. *Linguistics and Literary History: Essays in Stylistics* (Princeton Univ. Press, 1948).\n17. Chirosca, G., Rădvan, R., Muşat, S., Pop, M. & Chirosca, A. Machine learning models for artist classification of cultural heritage sketches. *Appl. Sci.***15** , 212 (2025).\n18. Theophilo, A., Padilha, R., Andaló, F. A. & Rocha, A. Explainable artificial intelligence for authorship attribution on social media. In *Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)* 2909–2913 (IEEE, 2022).\n19. Kong, Q., Choi, K. & Wang, Y. Large-scale MIDI-based composer classification. Preprint at [https://arxiv.org/abs/2010.14805](http://arxiv.org/abs/2010.14805)\n20. Kim, S., Lee, H., Park, S., Lee, J. & Choi, K. Deep composer classification using symbolic representation. In *Extended Abstracts for the Late-Breaking Demo Session of the 21st International Society for Music Information Retrieval Conference*[https://program.ismir2020.net/static/lbd/ISMIR2020-LBD-431-abstract.pdf](https://program.ismir2020.net/static/lbd/ISMIR2020-LBD-431-abstract.pdf) (2020).\n21. Tang, J., Wiggins, G. & Fazekas, G. Pianist identification using convolutional neural networks. In *Proc. 4th International Symposium on the Internet of Sounds* 1–6 (IEEE, 2023).\n22. Yang, D. & Tsai, T. Composer classification with cross-modal transfer learning and musically-informed augmentations. In *Proc. 22nd International Society for Music Information Retrieval Conference* (eds Lee, J. H. et al.) 802–809 (2021).\n23. Zhang, H., Karystinaios, E., Dixon, S., Widmer, G. & Cancino-Chacón, C. E. Symbolic music representations for classification tasks: a systematic evaluation. In *Proc. 24th International Society for Music Information Retrieval Conference* 848–858 (2023).\n24. Clayton, M., Rao, P., Shikarpur, N., Roychowdhury, S. & Li, J. Raga classification from vocal performances using multimodal analysis. In *Proc. 23rd International Society for Music Information Retrieval Conference* 283–290 (ISMIR, 2022).\n25. Mahmud Rafee, S. R., Fazekas, G. & Wiggins, G. HIPI: a hierarchical performer identification model based on symbolic representation of music. In *Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)* 1–5 (IEEE, 2023).\n26. Ramirez, R., Maestre, E. & Serra, X. Automatic performer identification in commercial monophonic jazz performances. *Pattern Recognit.***43** , 1514–1523 (2010).\n27. Falomir, Z., Museros, L., Sanz, I. & Gonzalez-Abril, L. Guessing art styles using qualitative colour descriptors, SVMs and logics. In *Artificial Intelligence Research and Development* 227–236 (IOS Press, 2009).\n28. Monson, I. T. *Saying Something: Jazz Improvisation and Interaction* (Univ. Chicago Press, 1996).\n29. Berliner, P. F. *Thinking in Jazz: The Infinite Art of Improvisation* (Chicago Univ. Press, 1994).\n30. Levine, M. *The Jazz Theory Book* (Sher Music, 1995).\n31. Kernfeld, B. *What to Listen for in Jazz* (Yale Univ. Press, 1995).\n32. Foscarin, F., Hoedt, K., Praher, V., Flexer, A. & Widmer, G. Concept-based techniques for ‘musicologist-friendly’ explanations in a deep music classifier. In *Proc. 23rd International Society for Music Information Retrieval Conference* (eds Rao, P. et al.) 876–883 (2022).\n33. Edwards, D., Dixon, S. & Benetos, E. PiJAMA: piano jazz with automatic MIDI annotations. *Trans. Int. Soc. Music Inf. Retr.***6** , 89–102 (2023).\n34. Rodríguez-Algarra, F., Sturm, B. L. & Dixon, S. Characterising confounding effects in music classification experiments through interventions. *Trans. Int. Soc. Music Inf. Retr.***2** , 52 (2019).\n35. Riley, X. & Dixon, S. Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline. In *Proc. 21st Sound and Music Computing Conference* 546–553 (2024).\n36. Kong, Q. et al. High-resolution piano transcription with pedals by regressing onset and offset times. *IEEE/ACM Trans. Audio Speech Lang. Process.***29** , 3707–3717 (2021).\n37. Breiman, L. Statistical modeling: the two cultures. *Stat. Sci.***16** , 199–215 (2001).\n38. Rudin, C. et al. Amazing things come from having many good models. In *Proc. 41st International Conference on Machine Learning* 42783–42795 (PMLR, 2024).\n39. Cheston, H. *Computational Modelling of Jazz Improvisation* . PhD thesis, Univ. Cambridge (2025).\n40. Breiman, L. Random forests. *Mach. Learn.***45** , 5–32 (2001).\n41. Frieler, K., Höger, F., Pfleiderer, M. & Dixon, S. Two web applications for exploring melodic patterns in jazz solos. In *Proc. 19th International Society for Music Information Retrieval Conference* (eds Gómez, E. et al.) 777–783 (2018).\n42. Conklin, D. Discovery of distinctive patterns in music. *Intell. Data Anal.***14** , 547–554 (2010).\n43. Naylor, J. *Jazz Vocab Book: Bill Evans. Jazz Vocab Series* (The Jazz Pursuit, 2024).\n44. Naylor, J. *Jazz Vocab Book: Oscar Peterson. Jazz Vocab Series* (The Jazz Pursuit, 2024).\n45. Levine, M. *The Jazz Piano Book* (Sher Music, 1989).\n46. Lyons, L. *The Great Jazz Pianists: Speaking of Their Lives and Music* (Quill, 1983).\n47. Ribeiro, M. T., Singh, S. & Guestrin, C. ‘Why should I trust you?’: explaining the predictions of any classifier. In *Proc.**22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD)* 1135–1144 (ACM, 2016).\n48. Zhang, H. & Dixon, S. Disentangling the Horowitz factor: learning content and style from expressive piano performance. In *Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)* 1–5 (IEEE, 2023).\n49. Ramoneda, P., Eremenko, V., D’Hooge, A., Parada-Cabaleiro, E. & Serra, X. Towards explainable and interpretable musical difficulty estimation: a parameter-efficient approach. In *Proc. 25th International Society for Music Information Retrieval Conference* 520–528 (ISMIR, 2024).\n50. Frieler, K., Pfleiderer, M., Zaddach, W.-G. & Abeßer, J. Midlevel analysis of monophonic jazz solos: a new approach to the study of improvisation. *J. New Music Res.***45** , 143–162 (2016).\n51. Weiß, C., Balke, S., Abeßer, J. & Müller, M. Computational corpus analysis: a case study on jazz solos. In *Proc. 19th International Society for Music Information Retrieval Conference* 416–423 (ISMIR, 2018).\n52. Benadon, F. Slicing the beat: jazz eighth-notes as expressive microrhythm. *Ethnomusicology***50** , 73–98 (2006).\n53. Gridley, M. C. *Jazz Styles: History and Analysis* 5th edn (Prentice Hall, 1997).\n54. Gioia, T. *The History of Jazz* 2nd edn (Oxford Univ. Press, 2011).\n55. Riley, X. & Dixon, S. Filobass: a dataset and corpus based study of jazz basslines. In *Proc. 24th International Society for Music Information Retrieval Conference* 500–507 (2023).\n56. Hamilton, M. & Pearce, M. Trajectories and revolutions in popular melody based on U.S. charts from 1950 to 2023. *Sci. Rep.***14** , 14749 (2024).\n57. Broze, Y. & Shanahan, D. Diachronic changes in jazz harmony: a cognitive perspective. *Music Percept.***31** , 32–45 (2013).\n58. Harrison, P. M. C. & Pearce, M. T. An energy-based generative sequence model for testing sensory theories of Western harmony. In *Proc. 19th International Society for Music Information Retrieval Conference* 160–167 (ISMIR, 2018).\n59. Cheston, H., Schlichting, J. L., Cross, I. & Harrison, P. M. C. Jazz trio database: automated annotation of jazz piano trio recordings processed using audio source separation. *Trans. Int. Soc. Music Inf. Retr.***7** , 144–158 (2024).\n60. Solovyev, R., Stempkovskiy, A. & Habruseva, T. Benchmarks and leaderboards for sound demixing tasks. Preprint at [https://arxiv.org/abs/2305.07489](http://arxiv.org/abs/2305.07489)\n61. Pfleiderer, M., Frieler, K., Abeßer, J., Zaddach, W.-G. & Burkhart, B. *Inside the Jazzomat: New Perspectives for Jazz Research* (Schott Campus, 2017).\n62. Alvarez, D. A. P., Gelbukh, A. & Sidorov, G. Composer classification using melodic combinatorial *n* -grams.*Expert Syst. Appl.***249** , 123300 (2024).\n63. Deepaisarn, S. et al. NLP-based music processing for composer classification. *Sci. Rep.***13** , 13228 (2023).\n64. Eppler, A., Mannchen, A., Abeßer, J., Weiß, C. & Frieler, K. Automatic style classification of jazz records with respect to rhythm, tempo, and tonality. In *Proc. 9th Conference on Interdisciplinary Musicology* (eds Klüche, T. & Miranda, E.) 162–167 (Fraunhofer Institute for Digital Media Technology, 2014).\n65. Saunders, C., Hardoon, D. R., Shawe-Taylor, J. & Widmer, G. Using string kernels to identify famous performers from their playing style. In *Machine Learning: ECML 2004* (eds Boulicaut, J.-F. et al.) 384–395 (Springer, 2004).\n66. Widmer, G. & Zanon, P. Automatic recognition of famous artists by machine. In *Proc. 16th European Conference on Artificial Intelligence* (eds López de Mántaras, R. & Saitta, L.) 1109–1110 (IOS Press, 2004).\n67. Pedregosa, F. et al. Scikit-learn: machine learning in Python. *J. Mach. Learn. Res.***12** , 2825–2830 (2011).\n68. Ens, J. & Pasquier, P. Quantifying musical style: ranking symbolic music based on similarity to a style. In *Proc. 20th International Society for Music Information Retrieval Conference* 870–877 (ISMIR, 2019).\n69. Sauvé, S. A. *Prediction in Polyphony: Modelling Musical Auditory Scene Analysis* . PhD thesis, Queen Mary Univ. London (2018).\n70. Chou, Y.-H., Chen, I.-C., Ching, J., Chang, C.-J. & Yang, Y.-H. MidiBERT-Piano: Large-scale pre-training for symbolic music classification tasks. *Journal of Creative Music Systems***8** (2024)\n71. Uitdenbogerd, A. L. & Zobel, J. Manipulation of music for melody matching. In *Proc. 6th ACM International Conference on Multimedia* 235–240 (Association for Computing Machinery, 1998).\n72. Norgaard, M., Bales, K. & Hansen, N. C. Linked auditory and motor patterns in the improvisation vocabulary of an artist-level jazz pianist. *Cognition***230** , 105293 (2023).\n73. Cross, P. & Goldman, A. Interval patterns are dependent on metrical position in jazz solos. *Music Percept.***39** , 299–312 (2022).\n74. Bantula, H., Giraldo, S. & Ramirez, R. Jazz ensemble expressive performance modeling. In *Proc. 17th International Society for Music Information Retrieval Conference* (eds Mandel, M. I. et al.)674–680 (ISMIR, 2016).\n75. Schonlau, M., Guenther, N. & Sucholutsky, I. Text mining with *n* -gram variables.*Stata J.***17** , 866–881 (2017).\n76. Bergstra, J., Bardenet, R., Bengio, Y. & Kégl, B. Algorithms for hyper-parameter optimization. In *Proc. 25th International Conference on Neural Information Processing Systems*[https://papers.nips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf](https://papers.nips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf) (2011).\n77. Cheston, H., Balen, J. V. & Durand, S. Automatic identification of samples in hip-hop music via multi-loss training and an artificial dataset. Preprint at [https://arxiv.org/abs/2502.06364](http://arxiv.org/abs/2502.06364) (2025).\n78. Paszke, A. et al. PyTorch: an imperative style, high-performance deep learning library. In *33rd Conference on Neural Information Processing Systems*[https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf](https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf) (2019).\n79. Flexer, A. & Schnitzer, D. Effects of album and artist filters in audio similarity computed for very large music databases. *Comput. Music J.***34** , 20–28 (2010).\n80. Mitchell, M. et al. Model cards for model reporting. In *Proc. Conference on Fairness, Accountability, and Transparency* 220–229 (Association for Computing Machinery, 2019).\n81. Cheston, H., Bance, R. & Harrison, P. Data from: Machine learning of artistic fingerprints in jazz. *Zenodo*[https://doi.org/10.5281/zenodo.14774191](https://doi.org/10.5281/zenodo.14774191) .(2025).\n82. Cheston, H. Code from: Machine learning of artistic fingerprints in jazz. *Zenodo*[https://doi.org/10.5281/zenodo.20345004](https://doi.org/10.5281/zenodo.20345004) (2026).\n83. Cheston, H. Aeberscale: a simple key-finding algorithm designed for jazz. *Zenodo*[https://doi.org/10.5281/zenodo.20345077](https://doi.org/10.5281/zenodo.20345077) (2026).\n84. Meredith, D. The ps13 pitch spelling algorithm. *J. New Music Res.***35** , 121–159 (2006).\n\n## Acknowledgements\n\nWe thank S. Dixon, M. Clayton, D. Tymoczko, B. Sober and an anonymous reviewer for their helpful comments on several initial drafts of this work. Parts of this research previously formed part of a PhD thesis by the first author, published<sup>[39](/articles/s42256-026-01279-9#ref-CR39)</sup> according to the requirements of the University of Cambridge.\n\n## Funding\n\nH.C. declares support for the research of this work from a PhD studentship from the Cambridge Trust (grant number 10615996). This study was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service ([www.csd3.cam.ac.uk](http://www.csd3.cam.ac.uk)), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (grant number EP/T022159/1), and DiRAC funding from the Science and Technology Facilities Council ([www.dirac.ac.uk](http://www.dirac.ac.uk)).\n\n## Ethics declarations\n\n### Competing interests\n\nThe authors declare no competing interests.\n\n## Peer review\n\n### Peer review information\n\n*Nature Machine Intelligence* thanks Barak Sober and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.\n\n## Additional information\n\n**Publisher’s note** Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.\n\n## Extended data\n\n### [Extended Data Fig. 1 Feature importance by domain and model type.](/articles/s42256-026-01279-9/figures/7)\n\nEach bar shows the mean loss in test accuracy from permuting all melody or harmony features (left panel) and bootstrapped subsamples of 2,000 features (right panel), stratified by model type. Data are presented as mean values with error bars showing ± 1 *SE* (with N = 1,000 bootstrap iterations). Corresponding data points are shown in the overlaid dot plots.\n\n### [Extended Data Fig. 2 Proposed multi-input architecture.](/articles/s42256-026-01279-9/figures/8)\n\nAn input transcription is represented using four piano rolls, each relating to a musical domain. Each roll is processed with a separate sub-network, the embeddings are then pooled, and class probabilities are generated. The architecture of an 18- layer ResNet is shown here; different architectures are tested in our experiments (Extended Data Table [3](/articles/s42256-026-01279-9#Tab3)). Optional components are represented with dashed lines.\n\n### [Extended Data Fig. 3 Class accuracy per musical domain.](/articles/s42256-026-01279-9/figures/9)\n\nEach facet shows the percentage of recordings in the held-out test split classified correctly for each pianist when predicting using only a single sub-network.\n\n### [Extended Data Fig. 4 Class accuracy across musical domains.](/articles/s42256-026-01279-9/figures/10)\n\nEach facet contains a heatmap showing the probability that, when given a recording and the output from a sub-network, the model will identify a particular pianist. The proportion of hits is shown for each facet along the diagonal; all other values are misses. Lighter and darker colours indicate lower and higher predictive probability, respectively.\n\n### [Extended Data Fig. 5 Melody extraction results.](/articles/s42256-026-01279-9/figures/11)\n\nThe top panel shows five seconds of a MIDI transcription of a performance by Kenny Barron, the middle panel shows the melody in the same transcription as annotated by the first author (a jazz expert), and the bottom panel shows the same transcription after applying the implementation of the ‘Skyline’ algorithm [71] used in this work. When the same transcription was processed using the model described by Chou et al. [70], no notes were detected as being part of the melody.\n\n### [Extended Data Fig. 6 LR accuracy with different values of n.](/articles/s42256-026-01279-9/figures/12)\n\nThe left panel shows changes in accuracy when the maximum value of n used to extract melody and harmony features increases, from 3 to 8 inclusive. The right panel shows equivalent changes when the minimum value of n increases. All accuracy scores are the results of optimising hyperparameters separately using random sampling for each value of n (Section 4.2.2 in the full text) and are obtained for the validation split of the dataset.\n\n### [Extended Data Fig. 7 Feature counts.](/articles/s42256-026-01279-9/figures/13)\n\nEach panel shows the counts of the twenty most frequently occurring (**a**) melody and (** b**) harmony features across all recordings and splits of the dataset. The x-axis values can be interpreted as the number of semitones from an initial starting note.\n\n## Supplementary information\n\n### [Supplementary Information (download PDF )](https://media.springernature.com/original/springer-static/esm/art%3A10.1038%2Fs42256-026-01279-9/MediaObjects/42256_2026_1279_MOESM1_ESM.pdf)\n\nSupplementary Sections 1–3, Figs. 1–56 and Tables 1–3.\n\n## Rights and permissions\n\n**Open Access**  This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit [http://creativecommons.org/licenses/by/4.0/](http://creativecommons.org/licenses/by/4.0/).\n\n## About this article\n\n### Cite this article\n\nCheston, H., Bance, R. & Harrison, P.M.C. Machine learning of artistic fingerprints in jazz.\n                    *Nat Mach Intell* **8**, 1261–1274 (2026). https://doi.org/10.1038/s42256-026-01279-9\n\n- Received:\n- Accepted:\n- Published:\n- Version of record:\n- Issue date:\n- DOI: https://doi.org/10.1038/s42256-026-01279-9", "url": "https://wpnews.pro/news/machine-learning-of-artistic-fingerprints-in-jazz", "canonical_source": "https://www.nature.com/articles/s42256-026-01279-9", "published_at": "2026-09-08 08:02:21+00:00", "updated_at": "2026-09-08 08:32:21.485600+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["Nature Machine Intelligence"], "alternates": {"html": "https://wpnews.pro/news/machine-learning-of-artistic-fingerprints-in-jazz", "markdown": "https://wpnews.pro/news/machine-learning-of-artistic-fingerprints-in-jazz.md", "text": "https://wpnews.pro/news/machine-learning-of-artistic-fingerprints-in-jazz.txt", "jsonld": "https://wpnews.pro/news/machine-learning-of-artistic-fingerprints-in-jazz.jsonld"}}