Probing the World for Groove
A comparison of a drum specific CNN and frozen PaSST transfer learning across 18,264 two bar grooves and 74 style labels.
Can a model trained on general audio understand drum style?
01 · Model work
Training and reshaping the models
Transformer modification
Froze the PaSST backbone, pretrained on AudioSet, and replaced both of its classifier heads with custom MLP heads: 2 to 7 layers, two bottleneck shapes, time patchout on and off.
Best: 4 layer head, 0.8752 macro-F1 · 21 PaSST runs
CNN architecture design
Built a VGG style CNN in PyTorch for drum audio: stacked 3 × 3 convolutions with max pooling, depth tuned across 5, 7 and 9 layers.
Best: 7 layers (16 to 192 channels), 0.908 macro-F1 · 13 CNN runs
Cloud GPU training
34 runs across 11 rounds on NVIDIA A100 and L4 GPUs in Google Colab, with CUDA device placement and PyTorch Lightning for the PaSST runs. Adam, batch 16, up to 50 epochs, early stopping.
15 runs on A100, 15 on L4 · 3 to 8 hours each
Inputs and augmentation
Two input pipelines from one set of clips: 16 kHz log mel spectrograms for the CNN, 32 kHz audio padded to 10 s for PaSST. Gaussian noise, room simulation and time stretch, plus three padding modes.
Augmentation lifted the CNN; reflection padding lifted PaSST
Every run with its configuration, GPU and notebook is in the experiment explorer; the full protocol is on Method.
02 · Presentation
Watch the talk
03 · At a glance
04 · The finding
A drum specific CNN trained from scratch reached the strongest aggregate score on the full 74 class task, yet frozen PaSST transfer learning won under data scarcity and organized drum style into a more stable, hierarchical representation.
Transfer learning helped, but its value depended on data availability, classifier design, and how well general AudioSet pretraining aligned with fine grained drum style distinctions. The story is not “one model beat the other.”
05 · Why drum style?
Not a single hit. Not a transcription. A whole groove.
Single hit classification
Is this one sound a snare, a kick, a cymbal? Timbre of an isolated onset.
Drum transcription
Notate every onset into MIDI: a symbolic event sequence, no style label.
Drum pattern style
Read an entire two bar performance and name its style: a higher level task, close to genre, that this thesis investigates directly from audio.
06 · The data
From professional grooves to a 74 class audio task
To avoid building a dataset from scratch, the project adopted the Groove MIDI Dataset (professional drummer performances with primary and secondary style annotations) and rendered 18,264 two bar audio clips.
- 1 GMD performances
- 2 two bar clips
- 3 16 kHz / 32 kHz mono
- 4 primary ⊕ secondary labels
- 5 74 style classes
- 6 CNN vs PaSST
- 7 macro-F1
Clips are mono, padded to a fixed length, split ≈80 / 10 / 10 with proportional class balancing. Concatenating primary and secondary annotations yields the 74 classes, an intentionally fine grained, imbalanced target (rock styles alone make up nearly two thirds of the data).
07 · Two ways to hear a groove
A task specific CNN versus a pretrained transformer
- Log mel spectrogram input at 16 kHz
- Stacked local convolutional filters
- VGG style, adapted for drums
- Best configuration: 7 convolutional layers
- Captures local spectrotemporal, drum specific cues
- Audio resampled to 32 kHz, padded to 10 s
- Patchout Audio Spectrogram Transformer
- Frozen backbone, pretrained on AudioSet
- Best head: 4 layer MLP classifier
- Brings broad general audio knowledge to the task
08 · The evidence
Two charts carry the argument
| Series | macro-F1 |
|---|---|
| CNN (exp 10.1) | 0.908 |
| PaSST (exp 11.2) | 0.8752 |
| Series | macro-F1 |
|---|---|
| PaSST (exp 3.2) | 0.3911 |
| CNN (exp 3.1) | 0.3267 |
The early rock only experiment scored a higher 0.9204, but on a narrower, easier label set, so it belongs to the experiment history, not the headline.
09 · What the models learned
Accuracy isn’t the whole story: the representations differ
Beyond F1, the two models organize drum style differently. PaSST shows tighter clustering by primary style and smaller mean centroid shifts under perturbation, evidence of a more hierarchical embedding. The CNN leans on local features and is more sensitive to class imbalance.
10 · Read with care
Limitations
- Labels derive from dataset style tags, with ambiguous, overlapping genre boundaries.
- The 74 class target is imbalanced; rock dominates.
- Two bar clips omit longer musical structure.
- PaSST was used frozen, not finetuned, which limits task adaptation.
- Results are specific to GMD and its rendering conditions.
- t-SNE is exploratory, not proof of semantic separation.
11 · TimelineWhat happened, in order
-
Release A research website for the thesis
An editorial site over the research: six pages, an experiment explorer, run comparison and an interactive t-SNE explorer, built from saved notebook outputs and the workbook.
Customer value. The work became explorable without opening a notebook, with every headline number traced to its source.
- Every experiment filterable, sortable and linked to its notebook
- The t-SNE explorer over the saved embeddings, at seven perplexities
- Citations, BibTeX and downloads in one place
-
Release Presented at ISMIR 2025
The Late-Breaking Demo, presented at the 26th ISMIR on the KAIST campus in Daejeon, South Korea, September 21–25.
Customer value. The thesis condensed for the music information retrieval community, with a paper, a poster and a recorded talk.
- A Late-Breaking Demo paper (CC BY 4.0) and an A0 poster
- A recorded presentation on YouTube
- Supported by an ISMIR First-time Authors Grant and Leiden University grants
-
Release Research code and thesis on GitHub
The notebooks behind the experiments, the thesis PDF and the presentation, in a public research repository.
Customer value. Anyone can read the experiments at the source, run by run.
- The experiment notebooks for the CNN and PaSST runs
- The final thesis PDF and presentation slides
-
Release The MSc thesis, completed
Probing the World for Groove, for the MSc Media Technology at LIACS, Leiden University: a drum specific CNN and frozen PaSST compared across 74 drum styles. Defended at Kunstinstituut Melly, Rotterdam.
Customer value. A controlled answer to whether general audio pretraining helps with drum style: it wins under data scarcity, while a drum specific CNN wins on the full task.
- 34 experiments across 11 rounds
- Full 74 class task: CNN 0.908 macro-F1, frozen PaSST 0.8752
- Low data: PaSST 0.3911 against the CNN at 0.3267
12 · ISMIR 2025 · Late-Breaking Demo
Read, cite, and watch the work
Daejeon, South Korea. The full thesis, the Late-Breaking Demo paper, the poster, and the recorded presentation are all one click away.