Can computers learn what makes the most iconic jazz musicians stand out?

Publicly released:
International
Photo by Dolo Iglesias on Unsplash
Photo by Dolo Iglesias on Unsplash

Music researchers have trained machine-learning models that can tell iconic jazz musicians apart from their unique playing styles - an approach they say could help researchers dive even deeper into what makes these performers stand out. The researchers say a trained ear can often recognise famous musicians from what they call 'fingerprints', specific melodies, chord progressions or rhythms that often appear in their work. Computers generally struggle to analyse audio recordings of jazz, so to make it easier, the researchers selected 84 hours of performance recordings from 20 famous jazz pianists, and converted them to a digital 'piano roll' which shows when notes are played and at what pitch. They then trained models on the performances, and the best-performing model was able to tell the musicians apart at 94.4% accuracy. The researchers say their model allows for computer analysis of melodies, harmonies, rhythms and volume, but there are still plenty of other elements that make up jazz music that still can't be fully captured.

News release

From: Springer Nature

Artificial intelligence: Revealing the fingerprints of jazz musicians

Machine learning models can identify jazz pianists from recordings and reveal the musical ‘fingerprints’ that make individual performers recognisable, according to a study investigating 20 famous jazz pianists. The research, published in Nature Machine Intelligence, could offer insight into artist attribution and style, cultural heritage, and music education.

Musicians can often be recognised by distinctive patterns, or ‘fingerprints’, in their work, which can include harmonic progressions, rhythmic structures, and melodic motifs. Jazz offers an opportunity to study these fingerprints because it brings together composition, improvisation, and performance that can be especially unique to a specific performer. However, computational analysis of jazz has been difficult because most performances exist as audio recordings, and models trained directly on audio can be hard for humans to interpret due to variables such as the environment of the recording and the types of equipment used to record.

Huw Cheston and colleagues trained supervised-learning models to identify performers from a curated dataset of 84 hours of recordings from 1,629 performances by 20 famous jazz pianists. These included recordings of Bill Evans, Oscar Peterson, Thelonious Monk, Chick Corea, Keith Jarrett, McCoy Tyner, and Ahmad Jamal. The recordings were converted into MIDI ‘piano roll’ format, which is a digital representation that shows when notes are played and at what pitch. The best-performing model identified performers with 94.4% accuracy, while a more interpretable model that separated melody, harmony, rhythm, and dynamics reached 91.3% accuracy. When tested individually, harmony gave the most accurate predictions, followed by rhythm and melody, while dynamics was the least accurate.

The authors suggest that their approach could help researchers explore the musical patterns that make performers distinct, including features that align with existing scholarly literature and others that have not previously been discussed. They have also released open-source models and a web application to explore the results. However, they note that musical dimensions such as melody and harmony can overlap, and that MIDI piano rolls cannot capture all aspects of jazz performance, including tonal and timbral qualities such as vibrato and pitch-bending. Future work could extend the approach to other genres, instruments, historical contexts, and less well-known or historically under-represented musicians.

Attachments

Note: Not all attachments are visible to the general public. Research URLs will go live after the embargo ends.

Research Springer Nature, Web page The URL will go live after the embargo ends
Journal/
conference:
Nature Machine Intelligence
Research:Paper
Organisation/s: University of Cambridge, UK
Funder: H.C. declares support for the research of this work from a PhD studentship from the Cambridge Trust (grant number 10615996). This study was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service (www.csd3.cam. ac.uk), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (grant number EP/T022159/1), and DiRAC funding from the Science and Technology Facilities Council (www.dirac.ac.uk).
Media Contact/s
Contact details are only visible to registered journalists.