← Probing the World for Groove

✳ Resources

Papers, data, notebooks and reproducibility

Everything needed to read, cite, inspect, or build on the work: the publication artifacts, the experiment workbook, and a curated index into the research notebooks.

01 · Downloads and links

Publication artifacts

02 · Hear the data

Two bar grooves across styles

A handful of rendered GMD clips, the same kind of two bar audio the models classify. Style and tempo come from the dataset annotations; no model predictions are shown.

Rock · shuffle 125 BPM · 3.8s
Funk 95 BPM · 5.1s
Jazz · fast 290 BPM · 1.7s
Hiphop 70 BPM · 6.9s
Latin · brazilian bossa 127 BPM · 3.8s
Afrobeat 94 BPM · 5.1s
Soul 98 BPM · 4.9s
Reggae 126 BPM · 3.8s

Excerpts from the Groove MIDI Dataset (Google / Magenta, CC BY 4.0), 16 kHz mono.

03 · Curated notebook index

Read the experiments at the source

Selected notebooks grouped by theme, each linking to the research repository. The Experiments explorer maps every one of the 34 experiments to its notebook.

Foundational and configuration

Dataset exploration (low data and full)

Model depth and patchout

Augmentation and padding

Representation analysis

Full archive: all notebooks on GitHub →

04 · Attribution and reproducibility

Groove MIDI Dataset

All audio derives from the Groove MIDI Dataset by Google / Magenta, released under CC BY 4.0. Any audio examples on this site are short excerpts used with attribution; the full dataset is not redistributed here.

Gillick, J., Roberts, A., Engel, J., Eck, D., & Bamman, D. (2019). Learning to Groove with Inverse Sequence Transformations. ICML.

How the work runs

  • Notebooks were developed and run in Google Colab on NVIDIA A100 or L4 GPUs (assigned by runtime availability); each experiment took roughly 3 to 8 hours depending on GPU type.
  • Key libraries: passt_hear21, DrumClassification (CNN basis), audiomentations.
  • This site builds from saved notebook outputs, and never retrains models or reexecutes notebooks.
  • Three notebooks were restored from local backups after surfacing as truncated in the repo (see the audit log).
05 · Bibliography

References cited in the thesis

The works cited in Probing the World for Groove, each linked to the paper, preprint, or repository.

  1. K. Koutini, J. Schlüter, H. Eghbal-zadeh, G. Widmer. Efficient Training of Audio Transformers with Patchout . arXiv:2110.05069, 2022.
  2. S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, F. Wei. BEATs: Audio Pre-Training with Acoustic Tokenizers . arXiv:2212.09058, 2022.
  3. K. Koutini, H. Eghbal-zadeh, G. Widmer. Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks . arXiv:2105.12395, 2021.
  4. A. Quelennec, M. Olvera, G. Peeters, S. Essid. On the Choice of the Optimal Temporal Support for Audio Classification with Pre-Trained Embeddings . arXiv:2312.14005, 2023.
  5. Y. Ding, A. Lerch. Audio Embeddings as Teachers for Music Classification . ISMIR · arXiv:2306.17424, 2023.
  6. K. Choi, G. Fazekas, M. Sandler, K. Cho. Convolutional Recurrent Neural Networks for Music Classification . arXiv:1609.04243, 2016.
  7. S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, et al.. CNN Architectures for Large-Scale Audio Classification . arXiv:1609.09430, 2017.
  8. J. Cramer, H.-H. Wu, J. Salamon, J. P. Bello. Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings . IEEE ICASSP, 2019.
  9. Y. Gong, Y.-A. Chung, J. Glass. AST: Audio Spectrogram Transformer . arXiv:2104.01778, 2021.
  10. M. Won, K. Choi, X. Serra. Semi-Supervised Music Tagging Transformer . arXiv:2111.13457, 2021.
  11. K. Hiner. Drum Classification . GitHub: khiner/DrumClassification, 2023.
  12. M. Yun, J. Bi. Deep Learning for Musical Instrument Recognition . Tech. Report, University of Rochester, 2018.
  13. R. Vogl. Deep Learning Methods for Drum Transcription and Drum Pattern Generation . Doctoral thesis, JKU Linz, 2018.
  14. L. Géré, N. Audebert, P. Rigaux. Improved Symbolic Drum Style Classification with Grammar-Based Hierarchical Representations . ISMIR · arXiv:2407.17536, 2024.
  15. J. Gillick, A. Roberts, J. Engel, D. Eck, D. Bamman. Learning to Groove with Inverse Sequence Transformations . ICML · arXiv:1905.06118, 2019.
  16. K. Choi, G. Fazekas, M. Sandler, K. Cho. Transfer Learning for Music Classification and Regression Tasks . ISMIR · arXiv:1703.09179, 2017.
  17. Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, M. Plumbley. PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition . arXiv:1912.10211, 2020.
  18. B. Elizalde, S. Deshmukh, M. Al Ismail, H. Wang. CLAP: Learning Audio Concepts from Natural Language Supervision . arXiv:2206.04769, 2022.
  19. H. Park, Y. Chung, J.-H. Kim. Deep Neural Networks-based Classification Methodologies of Speech, Audio and Music, and its Integration for Audio Metadata Tagging . Journal of Web Engineering 22(1), 2023.
  20. T. Morocutti, F. Schmid, K. Koutini, G. Widmer. Device-Robust Acoustic Scene Classification via Impulse Response Augmentation . arXiv:2305.07499, 2023.
  21. J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, et al.. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events . IEEE ICASSP, 2017.
  22. H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. Jégou. Training Data-Efficient Image Transformers & Distillation through Attention . arXiv:2012.12877, 2021.
  23. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, et al.. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale . arXiv:2010.11929, 2021.
  24. K. Koutini. PaSST . GitHub: kkoutini/PaSST, 2023.
  25. K. Koutini. PaSST hear21 . GitHub: kkoutini/passt_hear21, 2023.
  26. A. Gazneli, G. Zimmerman, T. Ridnik, G. Sharir, A. Noy. End-to-End Audio Strikes Back: Boosting Augmentations Towards an Efficient Audio Classification Network . arXiv:2204.11479, 2022.
  27. I. Jordal. audiomentations . GitHub: iver56/audiomentations, 2025.
  28. T. Eriksen. LUMT-Thesis-Resources . Google Drive, 2025.