HRNSXTN x RDMSXN
Team THIRI's project from a two podium weekend at the Music Hackspace × MUTEK hackathon in Montréal: a voice leading harmonizer that fits a stompbox, and a latent granular instrument where a folder of recordings becomes a playable sound world: whatever you play comes back one bar later, revoiced through that corpus. Both live on the Elk Stomp, a screenless embedded pedal.
A neural model too big for a device does not have to shrink; it has to think less often. Run the heavy model offline, distill its judgment into something tiny, and put decisions at 10 Hz so the audio thread only ever plays real recordings.
01 · Demo
See it running
02 · At a glance
- Role
- Audio ML (on a four person team), then solo on the embedded port
- Timeframe
- August 2026 to present
- Audience
- Performers who want a corpus of their own recordings under their hands with no laptop on stage, and embedded audio developers asking what ML honestly fits on hardware with no NPU.
- Collaborators
- Dennison Blackett, Radu-Alex Ceban, Jazz Calls Home
- Stack
-
- Elk Stomp (STM32MP157)
- Elk Audio OS / SUSHI
- JUCE 8 (VST3)
- PyTorch
- music2latent
- RTNeural
- gRPC / OSC / MIDI
The instrument was built around music2latent, a 58M parameter neural codec that measures slower than real time on a laptop CPU, and the target is a pedal with two 800 MHz cores, no neural accelerator, no PyTorch runtime for its architecture, and a hard 1.333 ms deadline on every audio callback. The team's hackathon plugin loaded on the board and averaged 6.8% of that budget, then peaked at 1302% and glitched: the deadline is per callback peak, not average. The port is therefore an architecture problem, not an optimization problem: which decisions genuinely need the network, and when.
Constraints that shaped it
- A hard realtime budget: 64 samples per callback at 48 kHz, 1.333 ms, on two Cortex-A7 cores without out of order execution
- No neural accelerator, and no PyTorch/LibTorch runtime exists for 32 bit armv7 at all
- No screen, no menus: the challenge brief was knobs, encoders, and a footswitch, played by feel
- Nothing may allocate, lock, or make a syscall on the audio thread (the dual kernel OS turns violations into audible glitches)
- The embedded C++ must reproduce the Python research instrument exactly; its selections are the sound
03 · Outcomes
What changed
- 1st place, Elk Audio Challenge; 2nd place, Roland Future Design Lab Challenge (Music Hackspace × MUTEK, Montréal 2026)
- Accepted to the ISMIR 2026 Late-Breaking Demo session (Abu Dhabi, November 2026); code, board logs and notebooks are public
- 58M parameter codec replaced at the input by a distilled 103k parameter student (560× smaller) that keeps its similarity ordering (Spearman ρ .83)
- Runs on the Elk Stomp itself, no laptop in the signal path: the student takes about 1% of each 93 ms analysis, and every analysis finished on time
- An 8 bit matching table made the corpus search 2.6 to 3.1 times faster with the ordering unchanged; golden parity held at 200/200 identical grain selections
- The team's posthackathon fix put five harmony voices on the pedal at 0.59× of the deadline (a ~50× per voice saving, and it follows the player's intonation)
- Two latent defects found in the platform's open source audio engine by reading its source
04 · Recognition
Montréal, then Abu Dhabi
Music Hackspace × MUTEK
Montréal, Canada · August 2026
1st place, Elk Audio Challenge; 2nd place, Roland Future Design Lab Challenge. Presented at the Winners Showcase, MUTEK Forum, August 28.
ISMIR 2026 · Late-Breaking Demo
Abu Dhabi, UAE · November 8–12, 2026
Accepted at the 27th ISMIR, with Jazz Segovia and Dennison Blackett. Code, board logs and notebooks are public.
“Neural Codec Latent Distillation for Corpus Retrieval Synthesis on the Elk Stomp”
A 103k parameter student distilled from the 58M parameter music2latent encoder keeps its similarity ordering (Spearman ρ .83) and runs on the pedal at about 1% of each 93 ms analysis.
Paper, code and measurements on GitHub05 · At MUTEK
On the official program
06 · Architecture
How it fits together
07 · The hardware
Knobs before screens
08 · The prototype
Three faces, on screen
Built and gated on a laptop the week before the hackathon, before it met the pedal.
09 · The storyHow it came together
The opportunity
The Elk Audio brief at MUTEK was blunt: no laptop, no menus, an instrument you work by feel. The Roland Future Design Lab brief asked for a neural model treated as something performable. Both are the same question from different sides: what does machine learning look like when it has to live in a musician’s hands rather than in a browser tab? Team THIRI entered both challenges and placed in both. The instrument on this page is my half of that answer (a corpus player where the intelligence chooses real recordings instead of synthesizing new ones), and the engineering that follows is what it took to make that honest on a pedal.
How it works, at the boundary
A folder of recordings is embedded once, offline, by music2latent into 64 dimensional frames at 10.7 Hz, and packed with its audio into a single memory mappable file. On the pedal, live input becomes mel frames; a distilled student network predicts where the big model would have placed them; cosine matching with temperature sampling picks corpus grains; and the audio thread plays those grains back (real recordings, windowed and crossfaded) one bar behind your playing. The latency is not hidden. It is the instrument’s character: you play, and the room answers.
What I chose, and why
Three calls, each a decision record below. First, the model was never compressed; the decisions moved: offline precomputation, a 560×-smaller distilled student, inference at the 10.7 Hz decision rate instead of the 48 kHz audio rate. Second, the neural decoder was deleted rather than shrunk: playback is real corpus audio, which keeps the sound accountable and cost nothing the ear could keep. Third, equivalence over elegance: the C++ port is held to bit identical grain selection against the Python reference, and the one divergence the parity harness caught was resolved in the reference’s favor, because its quirk is part of the sound.
Proven before the hardware, then on it
The riskiest question, whether cheap matching still feels like the big model’s taste, was answered before any hardware was touched. The finding that reframed it: at performance settings the original instrument agrees with itself on only 2.6% of picks across seeds, so the right measure is regret, and the student lands in the teacher’s top 4–7% of candidates (a raw MFCC baseline fails outright at 15–18%). In September the instrument went onto the pedal itself, every analysis stage was timed on the board, and the study became a Late-Breaking Demo accepted at ISMIR 2026. What remains open is a blind listening verdict. The prototyping before the hackathon (the studio interface with its live map of 17,000 grains) is in the research notes below.
10 · RoadmapNow, next, later
Present at ISMIR 2026
The Late-Breaking Demo session's deliverables: an A0 poster, a thumbnail and a captioned demo video of at most five minutes, for Abu Dhabi, November 8–12.
Blind listening verdict
The metrics say the student preserves the instrument's taste; ears decide. Nine blind coded clips per condition, identical playback everywhere, key held separately until the ranking is done.
A tiny decoder to win back morph
Mosaic playback deliberately gave up morph: sounds that exist between recordings. A DDSP style decoder fits the budget on paper (~10–20% of a core); whether it earns its place on textural corpora is a listening question, gated on the onboard numbers.
Desktop instrument end to end
The full chain (corpus pack, distilled student, exact parity matcher, one bar late grain playback) running under the platform's own audio engine with knob control.
First sound from the pedal
Cross built for the board's 32 bit ARM target and loaded under SUSHI without the vendor SDK; the answer bar now comes from the pedal's own output jacks.
Onboard ML benchmarks
Every analysis stage timed on the silicon: front end, student and corpus search at three corpus sizes, with 32 and 8 bit tables. These runs became the ISMIR 2026 paper's second experiment.
11 · TimelineWhat happened, in order
-
Release Paper, code and board logs published
A public repository behind the ISMIR 2026 paper: the desktop instrument's core, the pack builder, the distillation experiments, the board plugin and its cross build, and the board transcripts.
Customer value. Every number in the paper can be traced to code and reproduced from the shared data, and the board plugin can be rebuilt.
- Five notebooks, committed with outputs, each ending in the paper numbers it reproduces
- The trained student, GRU and ridge map weights (CC BY-NC 4.0)
- Pack format specification and selection parity tests
- Board scripts, SUSHI configurations and the Experiment 2 transcripts
-
Research · Experiment ISMIR 2026: does a distilled student choose as the codec would?
The Late-Breaking Demo accepted at ISMIR 2026 tests whether a 103k parameter student, distilled from music2latent, keeps the codec's similarity ordering, and what then limits real time operation on the Elk Stomp.
Insights
- Spearman ρ .829 for the student, against .797 for a ridge map and .575 for MFCCs, over 48 query and corpus pairs
- Its picks rank near the teacher's own second seed (percentile .945 against .981)
- On the pedal the student takes about 1% of each 93 ms analysis; only the search grows with the corpus, and an 8 bit table speeds it 2.6 to 3.1 times
-
Release Timed on the board, then made faster
Per stage timers on the pedal's worker, an 8 bit matching table, and a front end that no longer divides 64 bit integers.
Customer value. Bigger sound worlds fit the 93 ms analysis period, and the paper's on device numbers come straight from these runs.
- Front end: 16.7 → 4.0 ms per hop
- 8 bit table: search 2.6 to 3.1 times faster, ordering unchanged (ρ 1.000)
- 12,018 grain hop: 30.2 → 17.6 ms
- Golden parity held at 200/200; offline bounces byte identical
- Milestone Onboard ML benchmarks
-
Release Five sound worlds, switched live from the panel
The rotary picks between five corpora on the pedal, each loaded on a helper thread and swapped without touching the audio thread.
Customer value. One pedal, several instruments: a performer changes sound world midset, with no laptop.
- Live switching verified four times on the board, zero drops
- Matching spread across hops: board p99 340 → 95 ms, picks bit identical
- Two playing sessions on the board (10 and 15 minutes), clean
- World select on the rotary, output level on a pot
-
Release The instrument runs on the pedal
Cross built for the Elk Stomp's 32 bit ARM cores and loaded under SUSHI without the vendor SDK; the answer bar now comes from the pedal's own output jacks.
Customer value. No laptop on stage: the instrument the challenge asked for, played from the pedal by feel.
- Docker cross build: Ubuntu 24.04 armhf, GCC 13, JUCE 8
- Checksummed file transfer over the serial console, about 8 KB/s
- On the board: worker p50 93.3 ms, p99 94.9 ms, zero overruns, stalls or drops
- All five pot mappings and hold verified by hand
Risks
- Per stage costs on the board still unmeasured
- Milestone First sound from the pedal
-
Release The public talk at MUTEK, told straight
The team's Winners Showcase presentation for the Elk Audio Challenge (Grands Ballets, August 28): THIRI as a product, and the honest account of a desktop plugin meeting a hard real time pedal.
Customer value. Musicians hear "play one horn, hear five; it fits in a stompbox"; engineers get the measured path from 13× over deadline to five voices at 0.59× with headroom.
- THIRI as a family: standalone, VST, and the live demo at demo.thiri.ai
- The porting lesson, on the slide: know the specs before you rebuild
- Future directions: ML on device via distillation; no GPU on the fly model building
Risks
- Sponsor and festival marks are attribution, not endorsement
-
Release The instrument runs under Elk's engine
End to end on desktop, under the same engine the pedal runs: live input → student embedding → grain match over a 12,018 grain corpus → the answer, one bar late.
Customer value. The whole instrument is testable and audible before a single cross compile; configs and knob mappings transfer to hardware unchanged.
- Corpus pipeline: recordings → one pack file (18.6 min in 18 s, 84× real time)
- Golden parity: 200/200 identical grain selections vs the Python reference
- Async worker under the host: p99 cadence <99 ms, zero drops; audio thread 0.09%
- Ten knob mapped parameters; bar clock on the plugin's own sample counter
Risks
- A7 budgets are projected until the board benchmarks run
-
Research · Experiment Distillation judged by regret, not agreement
Can a 103k parameter student stand in for a 58M parameter codec at the instrument's front end? Four conditions, four corpora, held out queries, and a metric correction.
Insights
- The oracle agrees with itself on only 2.6% of picks at live settings; regret is the honest lens, not agreement
- The student lands in the oracle's top 4–7% (ρ .84); a raw MFCC baseline fails at 15–18%
- A recurrent variant tied the simple MLP everywhere; the stateless model shipped
-
Research · Synthesis What actually fits on a pedal with no NPU
An eight chapter, cited knowledge base of the Elk Stomp platform (hardware, OS, host internals, and the real budget for machine learning), verified claim by claim.
Insights
- The budget gap is categorical: ~2,000 MACs per sample at audio rate, millions per frame at 10.7 Hz
- No PyTorch runtime exists for the board's 32 bit architecture; inference compiles into the plugin
- Two latent host defects found by reading its source, both now design constraints
- Milestone Desktop instrument end to end
-
Incident · Sev2 Resolved Loaded fine, averaged 6.8%, peaked at 1302%
The team's plugin cross compiled, loaded on the Elk Stomp with all parameters live, then overloaded the 1.333 ms audio deadline the moment its DSP engaged.
Impact. No on device demo at the event. The board's telemetry: average 6.8% of budget, peak 1302%, roughly 13× over.
Detection. Audible glitching under load; confirmed postevent from the engine's per processor timing statistics recovered off the board.
Root cause. Frame based DSP built around 512 sample hops concentrates work into one callback in eight. The deadline is per callback peak, not average.
Fix. The team replaced the spectral shifter with time domain PSOLA (~50× cheaper per voice): five voices at 0.59× of the deadline. This instrument made the rule architecture: analysis on a 10.7 Hz worker, the audio thread only mixes, an allocation guard in debug.
Followups
- Fold "peak, not average" into the realtime review checklist (done)
- Verify no involuntary mode switches on the pedal (board session) (open)
Blameless note. Desktop plugin conventions assume block sizes this platform does not offer. The measurement became a design rule, a review gate, and a slide in the public talk.
-
Research · Experiment Prehackathon: a folder of recordings becomes an instrument
The week before the hackathon, the latent granular idea was built and gated on a laptop: encode a corpus, match live playing grain by grain in the latent space, answer a bar late.
Insights
- The latency became the identity: one block ahead reads as an answer, not a lag
- Antihub and stickiness penalties matter as much as similarity (worst corpus, 88% of picks from five grains)
- Five shaping controls were enough; presets mattered more than a sixth
12 · DecisionsDecision records
Accepted Play real grains instead of running a decoder
Context. The laptop instrument decodes latents back to audio through the codec: the most expensive stage, and the one a pedal with no NPU most clearly refuses.
Decision. Delete the decoder. Match in the learned space; play back int16 corpus audio with equal power crossfades from precomputed margins.
Why. The mosaic character (real grains, intelligently chosen) is the half audiences respond to, and it survives intact. Intelligence in the choosing, not the manufacturing.
Consequences. The audio thread only mixes and windows. Morph is a scoped out feature with a named path back, not a silent omission.
Accepted Move the decisions, not the model
Context. The instrument depends on a 58M parameter codec that is slower than real time even on a laptop CPU. The target is two 800 MHz cores, no NPU, and no PyTorch runtime exists for its 32 bit architecture at all.
Decision. Restructure. The codec embeds the corpus once, offline; a distilled 103k parameter student embeds live input at 10.7 Hz on a worker thread. The audio thread never runs a network.
Why. At audio rate the board affords ~2,000 multiply accumulates per sample; at the decision rate, millions per frame. Same silicon, three orders of magnitude apart.
Consequences. Student inference costs 1–2% of one core; the audio thread measures 0.09%. The quality question moved somewhere measurable (the regret study), and the recipe generalizes to any latency critical system with intermittent decisions.