← Probing the World for Groove

✳ Representations

What the learned embeddings reveal

Two models can score similarly yet organize drum style very differently. Beyond macro-F1, the embeddings expose how each model structures style, and how stable that structure is.

01 · Why look past accuracy

Macro-F1 says how often a model is right, not how it understands the task. To compare understanding, the high dimensional embeddings of each model are projected to two dimensions with t-SNE and summarized by class centroids, the average position of each style. Tracking how far those centroids move between training runs measures the stability of a model's representation.

02 · Embedding structure · interactive

Explore the embedding space

Switch between models, then move the slider to reproject the same embeddings at different t-SNE perplexities. Low perplexity emphasises local neighbourhoods; high perplexity favours global structure. PaSST keeps coherent clusters by primary style across the whole sweep, while the CNN organises into smaller, more local groups.

Drag to zoom, double click to reset, click a style in the legend to hide it. Move the slider to reproject the same embeddings at a different t-SNE perplexity.

03 · Where the models confuse styles

Confusion concentrates on neighbouring styles

Both models keep a strong diagonal but stumble on stylistically adjacent classes, like jazz fast with dance breakbeat and jazz mediumfast, and funk with hip hop. The CNN also overpredicts the dominant rock classes, a symptom of class imbalance.

74 by 74 confusion matrix for the final CNN, dominated by a strong diagonal with sparse off diagonal confusions.
Final CNN (exp 10.1): confusion across all 74 classes. · GMD_CNN_prototype6_4.ipynb
74 by 74 confusion matrix for the final PaSST model, with a strong diagonal and confusions among related styles.
Final PaSST (exp 11.2): confusion across all 74 classes. · PaSST_setup6_16.ipynb

04 · The caveat

05 · Stability under padding and augmentation

How far the centroids move, by intervention

Each bar is the mean distance a model's primary style centroids travelled from its baseline run under a given change. The scales are completely different: the CNN's space lurches under every intervention, reflection padding most of all, while PaSST barely moves. Read each model on its own axis.

CNN: centroid shift by intervention
Reflection padding 94.7
Circular padding 89.9
Time stretch 86.0
Noise + room 84.5
Noise, room + reflection 78.9
Mean primary style shift from the zero padding baseline. Reflection padding moves it most. · workbook · GMD_CNN_primary_embedding
PaSST: centroid shift by intervention
Time stretch 8.1
Deeper head 8.1
Circular padding 6.5
Mean primary style shift from the reflection padding baseline. Note the axis: about 11.7× smaller than the CNN. · workbook · PaSST_primary_embedding

And which styles move the most

CNN: most shifted primary styles
afrobeat 184
hiphop 142
blues 127
country 120
dance 114
funk 107
middleeastern 93.8
rock 90.3
punk 88.2
reggae 84.0
Under reflection padding, mean shift 94.66 · workbook · GMD_CNN_primary_embedding
PaSST: most shifted primary styles
punk 9.1
gospel 8.9
pop 8.9
soul 8.6
middleeastern 8.6
country 8.5
afrobeat 8.2
reggae 8.2
rock 8.2
blues 8.1
Mean shift 8.07 · workbook · PaSST_primary_embedding

06 · What it suggests

A more hierarchical space, at a cost in accuracy

Taken together, PaSST's small, evenly distributed shifts and tight primary style clustering suggest its AudioSet pretraining carries a stable, hierarchical “world model” of audio that transfers to drum style. The CNN's large, uneven shifts, concentrated in a handful of styles, point to a more local, task fitted representation that is more sensitive to class imbalance and to interventions like padding.

This is the heart of the thesis's nuance: the CNN wins on aggregate macro-F1, yet PaSST organizes style more robustly. Accuracy and representational quality are not the same axis.

← Back to the experiments