§ Glossary

Field terms, in plain language.

The vocabulary that recurs across the projects, defined for readers who don't live in audio ML or live performance tooling. Grouped by where each term shows up; crosslinks point to related terms.

Windchime

Voice, retrieval, and the audio corpus. · 20 terms
Artist authored corpus

The closed library of 406 handbuilt stems and field recordings that Windchime retrieves from, organized across eleven instrument roles. It grew in tracked batches (203, 292, 330, 369, then 406), each one a new index epoch. Its character is the piece's character, which is why it stays private.

See also Stem Instrument role Field recording

ASR (faster-whisper)

Automatic speech recognition, the step that turns a visitor's spoken words into text before they are embedded. Windchime runs faster-whisper on device, and no audio is recorded.

See also Audio language model (ALM) Embedding space

Audio language model (ALM)

A model that maps text and audio into one shared space, so a phrase can retrieve matching sound. Windchime treats the active ALM as an exchangeable variable behind a single embedding interface; the audit compares configurations of public ALMs (CLAP variants, CLaMP 3, and others), six of them deployable in the installation.

See also CLAP Embedding space Retrieval (not generation)

Bounded session

The installation's unit of interaction: a staged turn with an explicit start, a set number of spoken prompts, and an automatic reset, rather than an endless jam. The framing is what makes the piece legible to a first time visitor and safe to leave running.

See also Sound modes ASR (faster-whisper)

Catalog coverage

How much of the corpus a given ALM can actually reach across a balanced set of prompts. Choosing a model fixes a reachable region of the library before anyone speaks, which the audit studies as a property of model choice.

See also Audio language model (ALM) One per role selection

CLAP

Contrastive Language-Audio Pretraining. A family of audio language models that map text and audio into one shared space, so a spoken phrase can retrieve matching sound. Windchime's default is a CLAP model, used frozen as a feature extractor with no training.

See also Audio language model (ALM) Embedding space

Content aware windowing

Embedding each stem from the window where the music actually plays, rather than a fixed offset that might land on silence. It repairs retrieval vectors computed from a quiet or unrepresentative slice.

See also Stem Index epoch

Embedding space

The shared numerical space, learned by contrastive audio language pretraining, where a transcribed phrase and every stem become comparable vectors. Retrieval is finding the stems whose vectors sit nearest the phrase's vector.

See also Audio language model (ALM) CLAP FAISS

Failover chain

The planner's ladder of fallbacks: a hosted language model drops to a local model, which drops to a deterministic template. A persistent audio watchdog separately rescues the sound engine in place, so the piece keeps playing even fully offline.

See also Guarded planner

FAISS

A library for fast nearest neighbour search over vectors. Windchime ranks stems by cosine similarity to the spoken query using an exact FAISS index, one per ALM configuration.

See also Embedding space One per role selection

Field recording

Audio captured in the world rather than the studio, gathered while travelling. Field recordings sit in the corpus alongside musical stems and can be retrieved by the same voice query.

See also Stem Artist authored corpus

Guarded planner

The pattern where a language model fills a validated JSON schema instead of writing code directly; a compiler then renders safe output. It keeps generation inspectable and prevents the model from ever emitting raw code into a live installation.

See also Strudel Failover chain

Index epoch

A versioned record of each time the corpus is embedded or rebuilt, stored in the corpus database. Because retrieval results depend on which index produced them, every measurement is attributable to a specific epoch.

See also Content aware windowing Audio language model (ALM)

Instrument role

One of the eleven curatorial categories (bass, drums, keys, pad, field, and so on) that organize the corpus. The labels guide selection rather than assert ground truth, and they let the system guarantee variety by drawing at most one stem per role.

See also One per role selection Artist authored corpus

Live Muse

A distinct earlier installation of mine, shown at Mutaciones in Barcelona and different in nature from Windchime. Windchime built on it, adapting parts of its stem library and code, but the two are separate works.

See also Artist authored corpus

One per role selection

The policy that admits at most one stem per instrument role from the top ranked candidates. It guarantees timbral variety in the layered result instead of returning six near identical drum loops.

See also Instrument role FAISS Retrieval (not generation)

Retrieval (not generation)

Windchime answers a voice by finding real recordings that match it, rather than synthesizing new audio. Nothing in the audio path is generated, which keeps the output accountable to the artist's own material and the pipeline deterministic.

See also Audio language model (ALM) Artist authored corpus One per role selection

Sound modes

Three retrieval and playback presets. Soundscape layers overlapping decaying stems; Focused restricts retrieval to stems verified as audible and layers them tighter; Responsive also requires an early onset and plays one shots that fade to invite the next prompt.

See also Bounded session Retrieval (not generation)

Stem

A single short audio layer, one instrument or texture, that the installation can play alone or stack with others. Windchime's stems are roughly thirty second loops rendered from studio sessions or recorded in the field.

See also Artist authored corpus Field recording Content aware windowing

Strudel

A browser based live coding environment for music (a JavaScript cousin of TidalCycles). Windchime renders validated patterns into Strudel to produce sound in real time.

See also Guarded planner

HRNSXTN x RDMSXN

Grains, latents, and the pedal. · 7 terms
8 bit matching table

The corpus latents stored as 8 bit integers with one scale per row instead of 32 bit floats. On the pedal the search over every grain ran 2.6 to 3.1 times faster, with the similarity ordering unchanged.

See also Embedding space Latent granular resynthesis

Elk Audio OS

A low latency Linux for embedded audio, here on the Elk Stomp's two Cortex-A7 cores. Its host, SUSHI, runs the instrument as a plugin with an audio callback every 1.33 ms.

See also RTNeural

Granular synthesis

Building sound from very short slices ("grains") of recorded audio, windowed and overlapped. HRNSXTN's latent variant chooses which grains to play by matching in a learned embedding space, so the output is always real recorded material, intelligently selected, never synthesized from scratch.

See also Knowledge distillation Latent granular resynthesis Embedding space

Knowledge distillation

Training a small "student" model to imitate a large "teacher" model's outputs rather than learning from raw labels. In HRNSXTN, a 103k parameter student learns to predict where a 58M parameter audio codec would place live sound in its embedding space: 560× smaller, cheap enough for a pedal with no neural accelerator.

See also Embedding space Granular synthesis music2latent

Latent granular resynthesis

Encoding an input and a corpus with a neural audio codec, replacing each input frame with a similar corpus frame, and decoding the result. On the pedal HRNSXTN keeps the matching but plays stored grains instead of decoding.

See also Granular synthesis music2latent

music2latent

A neural audio codec that turns about 93 ms of audio into one 64 dimensional latent and back. At 58M parameters it cannot run on the Elk Stomp, so HRNSXTN encodes the corpus offline and distills its encoder for the live input.

See also Knowledge distillation Latent granular resynthesis Embedding space

RTNeural

A small C++ library for running neural networks inside realtime audio code. It runs HRNSXTN's distilled student on the pedal in about 1% of each 93 ms analysis.

See also Knowledge distillation Elk Audio OS

Lichtspiel

Live performance, generation, and the set. · 11 terms
Capability adaptive mapping

The control layer that detects which monome is plugged in and folds each sketch onto it, from a small Grid and two encoder Arc up to a large Grid and four encoder Arc. The show plays on whatever hardware survives the trip, or none at all.

See also monome Digital twin

Idiom

A reusable control building block a visual scene is composed from: a fader bank, arc macros, a step sequencer, a cell painter. Idioms give every generated scene a consistent, playable relationship to the monome.

See also monome Visual parameter vector

LISTEN mode

The hands free performance mode where the instrument plays itself by following the live set, its sections, builds, and accents, rather than a fixed clock. It is one of three postures (Manual, Auto, LISTEN); the animated digital twin on these pages is showing LISTEN.

See also Digital twin Capability adaptive mapping

Max for Live

Ableton Live's embedded visual programming environment (Max/MSP). Lichtspiel uses a thin Max device as its bridge into the running set, reading clips, scenes, and transport.

See also Node bridge MIR pipeline

MIR pipeline

Music information retrieval: extracting musical features such as tempo, key, energy, and section structure from audio. Lichtspiel derives these from the running Ableton set and uses them, with a text prompt, to author visuals.

See also Max for Live Node bridge

Node bridge

The service between Ableton and the visuals, normalizing set state from the Max device and routing it to the p5 runtime and the authoring time model service. It is also where generated output is validated before it can reach the stage.

See also Max for Live Runtime purity

Runtime purity

Lichtspiel's rule that the live performance path never calls a model or the network. All AI happens at authoring time, so the visuals never stutter waiting on anything, even offline.

See also Node bridge Validation gate TouchDesigner

Self repair loop

The bounded retry around generation: when a scene fails a validation gate, the error is fed back and the scene is regenerated, up to a fixed number of attempts. It turns most near misses into shippable scenes with no human in the loop.

See also Validation gate

TouchDesigner

A popular node based visual programming tool for realtime graphics. It is powerful but demands learning a separate node graph craft; Lichtspiel's premise is that a musician should get set aware visuals without adopting a tool like it.

See also Runtime purity

Validation gate

One check in the five gate chain every generated scene must pass before it can play: strict typechecking, an allowlist lint, a monome playability marker check, and a headless render smoke test. A bounded self repair loop retries a failing scene before giving up.

See also Self repair loop Runtime purity

Visual parameter vector

The single shared set of numbers that describes a scene's look at any moment. Fixing it early let hardware, keyboard, and generated code all drive the same scenes through one contract, and it is the seam Windchime and Lichtspiel share.

See also Idiom

Probing the World for Groove

Transfer learning, evaluation, and drum style. · 7 terms
Centroid shift

How far a style's average position in a model's embedding space moves between two training runs. Small, even shifts suggest a stable representation; shifts compare only within one model.

See also t-SNE Transfer learning

Data augmentation

Training on altered copies of the audio (added noise, room simulation, time stretch) so a model learns what should not matter. In the thesis it lifted the CNN to its best score and lowered PaSST's.

See also macro-F1

Groove MIDI Dataset

Google Magenta's performances by professional drummers with style labels, released under CC BY 4.0. The thesis rendered 18,264 two bar clips from it, across 74 style classes.

See also macro-F1 Data augmentation

macro-F1

The F1 score computed for each class and then averaged, so each of the 74 drum styles counts equally however rare it is. It is the thesis's primary metric.

See also Transfer learning Groove MIDI Dataset

PaSST

The Patchout Audio Spectrogram Transformer, pretrained on AudioSet. The thesis froze it, replaced its classifier heads with small MLPs, and compared it with a CNN trained from scratch.

See also Transfer learning macro-F1

t-SNE

A projection that lays high dimensional embeddings out on a 2D map so that near neighbours stay near. Good for seeing how a model groups styles; distances on the map are not literal.

See also Centroid shift

Transfer learning

Reusing a model trained on one task as the starting point for another. The thesis keeps PaSST's AudioSet knowledge frozen and trains only a small classifier on top, to see what general audio pretraining already knows about drum style.

See also PaSST macro-F1

Shared

Hardware, reliability, and process. · 8 terms
Blameless postmortem

A writeup of an incident that focuses on the systemic cause and the fix rather than assigning fault. The goal is a lesson the system keeps, so the same failure does not recur.

See also Soak test

Decision record (ADR)

A short written record of a judgment call: the context, the options weighed, the decision, and its consequences. Capturing the reasoning at the moment of the choice is what lets it survive after the fact.

Digital twin

An onscreen mirror of the connected monome that shows every LED live and takes clicks and drags as input. It makes the hardware mapping legible before anyone touches it, and keeps both pieces playable with no hardware at all.

See also monome Varibright LISTEN mode

monome

A minimalist grid and encoder hardware controller: a Grid of backlit buttons and an Arc of rotary encoders. Windchime and Lichtspiel treat it as an expressive instrument, not a bank of switches.

See also serialosc Digital twin Varibright

Redaction gate

An automated check that scans anything derived from private repos for secrets, hostnames, emails, and large code blocks before it can ship. It is the firewall that lets this public site draw on private source without leaking it.

serialosc

The small background service that connects monome hardware to software over OSC, so a browser or app can read button presses and light the LEDs.

See also monome

Soak test

A long duration reliability run (hours, not minutes) that exercises a system continuously to surface slow leaks, timing drift, and rare failures a short test would never hit.

See also Blameless postmortem

Varibright

A monome grid whose buttons light at sixteen brightness levels rather than just on or off. Windchime and Lichtspiel use the extra range to show state and motion in the LEDs, not only which key is pressed.

See also monome Digital twin