← All projects
Transfer learning for drum audio style classification Published

Probing the World for Groove

A comparison of a drum specific CNN and frozen PaSST transfer learning across 18,264 two bar grooves and 74 style labels.

✳ The question

Can a model trained on general audio understand drum style?

ISMIR 2025 Late-Breaking Demo, KAIST, Daejeon; MSc thesis, LIACS, Leiden University, defended at Kunstinstituut Melly, Rotterdam PythonPyTorchPyTorch LightningCUDAA100 / L4 GPUs (Google Colab)PaSST (hear21passt)torchaudioaudiomentationsTensorFlow Datasetsscikit-learnpandasNumPymatplotlib / seabornPlotly
ISMIR 2025, Daejeon KAIST ISMIR, International Society for Music Information Retrieval Universiteit Leiden LIACS, Leiden Institute of Advanced Computer Science Kunstinstituut Melly
Two t-SNE maps of the same two bar drum clips, coloured by primary style: the CNN on the left, frozen PaSST on the right

01 · Model work

Training and reshaping the models

PyTorch PyTorch Lightning CUDA Cloud GPU training (A100, L4) Transformer modification (PaSST heads) CNN architecture design Transfer learning Audio augmentation
PaSST

Transformer modification

Froze the PaSST backbone, pretrained on AudioSet, and replaced both of its classifier heads with custom MLP heads: 2 to 7 layers, two bottleneck shapes, time patchout on and off.

Best: 4 layer head, 0.8752 macro-F1 · 21 PaSST runs

CNN

CNN architecture design

Built a VGG style CNN in PyTorch for drum audio: stacked 3 × 3 convolutions with max pooling, depth tuned across 5, 7 and 9 layers.

Best: 7 layers (16 to 192 channels), 0.908 macro-F1 · 13 CNN runs

CUDA

Cloud GPU training

34 runs across 11 rounds on NVIDIA A100 and L4 GPUs in Google Colab, with CUDA device placement and PyTorch Lightning for the PaSST runs. Adam, batch 16, up to 50 epochs, early stopping.

15 runs on A100, 15 on L4 · 3 to 8 hours each

torchaudio

Inputs and augmentation

Two input pipelines from one set of clips: 16 kHz log mel spectrograms for the CNN, 32 kHz audio padded to 10 s for PaSST. Gaussian noise, room simulation and time stretch, plus three padding modes.

Augmentation lifted the CNN; reflection padding lifted PaSST

Every run with its configuration, GPU and notebook is in the experiment explorer; the full protocol is on Method.

02 · Presentation

Watch the talk

03 · At a glance

18,264
two bar clips
Groove MIDI Dataset, rendered to audio
74
style classes
primary ⊕ secondary annotations
34
experiments
across 11 rounds
11
rounds
in 4 thematic categories

04 · The finding

A drum specific CNN trained from scratch reached the strongest aggregate score on the full 74 class task, yet frozen PaSST transfer learning won under data scarcity and organized drum style into a more stable, hierarchical representation.

Transfer learning helped, but its value depended on data availability, classifier design, and how well general AudioSet pretraining aligned with fine grained drum style distinctions. The story is not “one model beat the other.”

05 · Why drum style?

Not a single hit. Not a transcription. A whole groove.

Single hit classification

Is this one sound a snare, a kick, a cymbal? Timbre of an isolated onset.

Drum transcription

Notate every onset into MIDI: a symbolic event sequence, no style label.

Drum pattern style

Read an entire two bar performance and name its style: a higher level task, close to genre, that this thesis investigates directly from audio.

06 · The data

From professional grooves to a 74 class audio task

To avoid building a dataset from scratch, the project adopted the Groove MIDI Dataset (professional drummer performances with primary and secondary style annotations) and rendered 18,264 two bar audio clips.

  1. 1 GMD performances
  2. 2 two bar clips
  3. 3 16 kHz / 32 kHz mono
  4. 4 primary ⊕ secondary labels
  5. 5 74 style classes
  6. 6 CNN vs PaSST
  7. 7 macro-F1

Clips are mono, padded to a fixed length, split ≈80 / 10 / 10 with proportional class balancing. Concatenating primary and secondary annotations yields the 74 classes, an intentionally fine grained, imbalanced target (rock styles alone make up nearly two thirds of the data).

07 · Two ways to hear a groove

A task specific CNN versus a pretrained transformer

CNN learned from scratch
  • Log mel spectrogram input at 16 kHz
  • Stacked local convolutional filters
  • VGG style, adapted for drums
  • Best configuration: 7 convolutional layers
  • Captures local spectrotemporal, drum specific cues
PaSST transfer learning
  • Audio resampled to 32 kHz, padded to 10 s
  • Patchout Audio Spectrogram Transformer
  • Frozen backbone, pretrained on AudioSet
  • Best head: 4 layer MLP classifier
  • Brings broad general audio knowledge to the task

08 · The evidence

Two charts carry the argument

On the full 74 class task, the CNN edges ahead
On the full 74 class task, the CNN edges ahead 0 0.25 0.5 0.75 1 CNN exp 10.1 0.908 PaSST exp 11.2 0.8752
On the full 74 class task, the CNN edges ahead (macro-F1)
Seriesmacro-F1
CNN (exp 10.1) 0.908
PaSST (exp 11.2) 0.8752
Best configuration per model, GMD-full, 74 classes, a directly comparable pair. · Source: Thesis Experiment results.xlsx · Overall
But under data scarcity, the ranking reverses
But under data scarcity, the ranking reverses 0 0.25 0.5 0.75 1 PaSST exp 3.2 0.3911 CNN exp 3.1 0.3267
But under data scarcity, the ranking reverses (macro-F1)
Seriesmacro-F1
PaSST (exp 3.2) 0.3911
CNN (exp 3.1) 0.3267
GMD-mini (≈10% subset): frozen transfer features give PaSST the advantage when data is limited. · Source: Thesis Experiment results.xlsx · Dataset Exploration

The early rock only experiment scored a higher 0.9204, but on a narrower, easier label set, so it belongs to the experiment history, not the headline.

09 · What the models learned

Accuracy isn’t the whole story: the representations differ

Beyond F1, the two models organize drum style differently. PaSST shows tighter clustering by primary style and smaller mean centroid shifts under perturbation, evidence of a more hierarchical embedding. The CNN leans on local features and is more sensitive to class imbalance.

See the representation analysis →

10 · Read with care

Limitations

  • Labels derive from dataset style tags, with ambiguous, overlapping genre boundaries.
  • The 74 class target is imbalanced; rock dominates.
  • Two bar clips omit longer musical structure.
  • PaSST was used frozen, not finetuned, which limits task adaptation.
  • Results are specific to GMD and its rendering conditions.
  • t-SNE is exploratory, not proof of semantic separation.
11 · Timeline

What happened, in order

  1. Release A research website for the thesis

    An editorial site over the research: six pages, an experiment explorer, run comparison and an interactive t-SNE explorer, built from saved notebook outputs and the workbook.

    Customer value. The work became explorable without opening a notebook, with every headline number traced to its source.

    • Every experiment filterable, sortable and linked to its notebook
    • The t-SNE explorer over the saved embeddings, at seven perplexities
    • Citations, BibTeX and downloads in one place
  2. Release Presented at ISMIR 2025

    The Late-Breaking Demo, presented at the 26th ISMIR on the KAIST campus in Daejeon, South Korea, September 21–25.

    Customer value. The thesis condensed for the music information retrieval community, with a paper, a poster and a recorded talk.

    • A Late-Breaking Demo paper (CC BY 4.0) and an A0 poster
    • A recorded presentation on YouTube
    • Supported by an ISMIR First-time Authors Grant and Leiden University grants
  3. Release Research code and thesis on GitHub

    The notebooks behind the experiments, the thesis PDF and the presentation, in a public research repository.

    Customer value. Anyone can read the experiments at the source, run by run.

    • The experiment notebooks for the CNN and PaSST runs
    • The final thesis PDF and presentation slides
  4. Release The MSc thesis, completed

    Probing the World for Groove, for the MSc Media Technology at LIACS, Leiden University: a drum specific CNN and frozen PaSST compared across 74 drum styles. Defended at Kunstinstituut Melly, Rotterdam.

    Customer value. A controlled answer to whether general audio pretraining helps with drum style: it wins under data scarcity, while a drum specific CNN wins on the full task.

    • 34 experiments across 11 rounds
    • Full 74 class task: CNN 0.908 macro-F1, frozen PaSST 0.8752
    • Low data: PaSST 0.3911 against the CNN at 0.3267

12 · ISMIR 2025 · Late-Breaking Demo

Read, cite, and watch the work

Daejeon, South Korea. The full thesis, the Late-Breaking Demo paper, the poster, and the recorded presentation are all one click away.