← All projects
Neural audio instruments for a screenless pedal Active

HRNSXTN x RDMSXN

Team THIRI's project from a two podium weekend at the Music Hackspace × MUTEK hackathon in Montréal: a voice leading harmonizer that fits a stompbox, and a latent granular instrument where a folder of recordings becomes a playable sound world: whatever you play comes back one bar later, revoiced through that corpus. Both live on the Elk Stomp, a screenless embedded pedal.

✳ Thesis

A neural model too big for a device does not have to shrink; it has to think less often. Run the heavy model offline, distill its judgment into something tiny, and put decisions at 10 Hz so the audio thread only ever plays real recordings.

Music Hackspace × MUTEK hackathon: 1st, Elk Audio Challenge; 2nd, Roland Future Design Lab; presented at MUTEK Forum, Montréal 2026; ISMIR 2026 Late-Breaking Demo, Abu Dhabi Elk StompElk Audio OS / SUSHIC++ / JUCE 8RTNeuralARM NEONPyTorchmusic2latentDocker cross build
MUTEK, international festival of digital creativity and electronic music Elk Audio Roland Future Design Lab Music Hackspace ISMIR 2026, Abu Dhabi ISMIR, International Society for Music Information Retrieval
The Elk Stomp development board on a dark desk beside an iMac running the desktop devkit

01 · Demo

See it running

02 · At a glance

Role
Audio ML (on a four person team), then solo on the embedded port
Timeframe
August 2026 to present
Audience
Performers who want a corpus of their own recordings under their hands with no laptop on stage, and embedded audio developers asking what ML honestly fits on hardware with no NPU.
Collaborators
Dennison Blackett, Radu-Alex Ceban, Jazz Calls Home
Stack
  • Elk Stomp (STM32MP157)
  • Elk Audio OS / SUSHI
  • JUCE 8 (VST3)
  • PyTorch
  • music2latent
  • RTNeural
  • gRPC / OSC / MIDI
The problem

The instrument was built around music2latent, a 58M parameter neural codec that measures slower than real time on a laptop CPU, and the target is a pedal with two 800 MHz cores, no neural accelerator, no PyTorch runtime for its architecture, and a hard 1.333 ms deadline on every audio callback. The team's hackathon plugin loaded on the board and averaged 6.8% of that budget, then peaked at 1302% and glitched: the deadline is per callback peak, not average. The port is therefore an architecture problem, not an optimization problem: which decisions genuinely need the network, and when.

Constraints that shaped it 5
  • A hard realtime budget: 64 samples per callback at 48 kHz, 1.333 ms, on two Cortex-A7 cores without out of order execution
  • No neural accelerator, and no PyTorch/LibTorch runtime exists for 32 bit armv7 at all
  • No screen, no menus: the challenge brief was knobs, encoders, and a footswitch, played by feel
  • Nothing may allocate, lock, or make a syscall on the audio thread (the dual kernel OS turns violations into audible glitches)
  • The embedded C++ must reproduce the Python research instrument exactly; its selections are the sound

03 · Outcomes

What changed

  • 1st place, Elk Audio Challenge; 2nd place, Roland Future Design Lab Challenge (Music Hackspace × MUTEK, Montréal 2026)
  • Accepted to the ISMIR 2026 Late-Breaking Demo session (Abu Dhabi, November 2026); code, board logs and notebooks are public
  • 58M parameter codec replaced at the input by a distilled 103k parameter student (560× smaller) that keeps its similarity ordering (Spearman ρ .83)
  • Runs on the Elk Stomp itself, no laptop in the signal path: the student takes about 1% of each 93 ms analysis, and every analysis finished on time
  • An 8 bit matching table made the corpus search 2.6 to 3.1 times faster with the ordering unchanged; golden parity held at 200/200 identical grain selections
  • The team's posthackathon fix put five harmony voices on the pedal at 0.59× of the deadline (a ~50× per voice saving, and it follows the player's intonation)
  • Two latent defects found in the platform's open source audio engine by reading its source

04 · Recognition

Montréal, then Abu Dhabi

Music Hackspace × MUTEK

Montréal, Canada · August 2026

1st place, Elk Audio Challenge; 2nd place, Roland Future Design Lab Challenge. Presented at the Winners Showcase, MUTEK Forum, August 28.

ISMIR 2026 · Late-Breaking Demo

Abu Dhabi, UAE · November 8–12, 2026

Accepted at the 27th ISMIR, with Jazz Segovia and Dennison Blackett. Code, board logs and notebooks are public.

“Neural Codec Latent Distillation for Corpus Retrieval Synthesis on the Elk Stomp”

A 103k parameter student distilled from the 58M parameter music2latent encoder keeps its similarity ordering (Spearman ρ .83) and runs on the pedal at about 1% of each 93 ms analysis.

Paper, code and measurements on GitHub

05 · At MUTEK

On the official program

06 · Architecture

How it fits together

Corpus a folder of recordings
music2latent offline, on a laptop
Pack file grains + embeddings
Live match student MLP, 10.7 Hz
Grain playback real audio, 48 kHz
Knobs Sensei / Guru
The answer one bar later
The heavy model runs once, offline. On the pedal, decisions happen at 10.7 Hz and the audio thread only plays real recordings.

07 · The hardware

Knobs before screens

The Elk Stomp development board: eight pots, four encoders, footswitches and an OLED on a dark ground
The Elk Stomp development board is the whole interface: eight pots, four endless encoders, four footswitches and a small OLED that the instrument deliberately ignores. Shaping controls, grain size, wet and randomness sit on the pots and the rotary switches between five sound worlds live, so the corpus is played by feel, the way the challenge asked.

08 · The prototype

Three faces, on screen

09 · The story

How it came together

The opportunity

The Elk Audio brief at MUTEK was blunt: no laptop, no menus, an instrument you work by feel. The Roland Future Design Lab brief asked for a neural model treated as something performable. Both are the same question from different sides: what does machine learning look like when it has to live in a musician’s hands rather than in a browser tab? Team THIRI entered both challenges and placed in both. The instrument on this page is my half of that answer (a corpus player where the intelligence chooses real recordings instead of synthesizing new ones), and the engineering that follows is what it took to make that honest on a pedal.

How it works, at the boundary

A folder of recordings is embedded once, offline, by music2latent into 64 dimensional frames at 10.7 Hz, and packed with its audio into a single memory mappable file. On the pedal, live input becomes mel frames; a distilled student network predicts where the big model would have placed them; cosine matching with temperature sampling picks corpus grains; and the audio thread plays those grains back (real recordings, windowed and crossfaded) one bar behind your playing. The latency is not hidden. It is the instrument’s character: you play, and the room answers.

What I chose, and why

Three calls, each a decision record below. First, the model was never compressed; the decisions moved: offline precomputation, a 560×-smaller distilled student, inference at the 10.7 Hz decision rate instead of the 48 kHz audio rate. Second, the neural decoder was deleted rather than shrunk: playback is real corpus audio, which keeps the sound accountable and cost nothing the ear could keep. Third, equivalence over elegance: the C++ port is held to bit identical grain selection against the Python reference, and the one divergence the parity harness caught was resolved in the reference’s favor, because its quirk is part of the sound.

Proven before the hardware, then on it

The riskiest question, whether cheap matching still feels like the big model’s taste, was answered before any hardware was touched. The finding that reframed it: at performance settings the original instrument agrees with itself on only 2.6% of picks across seeds, so the right measure is regret, and the student lands in the teacher’s top 4–7% of candidates (a raw MFCC baseline fails outright at 15–18%). In September the instrument went onto the pedal itself, every analysis stage was timed on the board, and the study became a Late-Breaking Demo accepted at ISMIR 2026. What remains open is a blind listening verdict. The prototyping before the hackathon (the studio interface with its live map of 17,000 grains) is in the research notes below.

10 · Roadmap

Now, next, later

Now 1
Research in the open

Present at ISMIR 2026

The Late-Breaking Demo session's deliverables: an A0 poster, a thumbnail and a captioned demo video of at most five minutes, for Abu Dhabi, November 8–12.

Planned confidence: high
Next 1
Honest evaluation

Blind listening verdict

The metrics say the student preserves the instrument's taste; ears decide. Nine blind coded clips per condition, identical playback everywhere, key held separately until the ranking is done.

Planned confidence: high
Later 1
Scope out loud

A tiny decoder to win back morph

Mosaic playback deliberately gave up morph: sounds that exist between recordings. A DDSP style decoder fits the budget on paper (~10–20% of a core); whether it earns its place on textural corpora is a listening question, gated on the onboard numbers.

Planned confidence: low
Shipped 3
Make it real before the hardware

Desktop instrument end to end

The full chain (corpus pack, distilled student, exact parity matcher, one bar late grain playback) running under the platform's own audio engine with knob control.

Shipped confidence: high
No laptop on stage

First sound from the pedal

Cross built for the board's 32 bit ARM target and loaded under SUSHI without the vendor SDK; the answer bar now comes from the pedal's own output jacks.

Shipped confidence: high
Numbers over claims

Onboard ML benchmarks

Every analysis stage timed on the silicon: front end, student and corpus search at three corpus sizes, with 32 and 8 bit tables. These runs became the ISMIR 2026 paper's second experiment.

Shipped confidence: high
11 · Timeline

What happened, in order

  1. Release Paper, code and board logs published

    A public repository behind the ISMIR 2026 paper: the desktop instrument's core, the pack builder, the distillation experiments, the board plugin and its cross build, and the board transcripts.

    Customer value. Every number in the paper can be traced to code and reproduced from the shared data, and the board plugin can be rebuilt.

    • Five notebooks, committed with outputs, each ending in the paper numbers it reproduces
    • The trained student, GRU and ridge map weights (CC BY-NC 4.0)
    • Pack format specification and selection parity tests
    • Board scripts, SUSHI configurations and the Experiment 2 transcripts
  2. Research · Experiment ISMIR 2026: does a distilled student choose as the codec would?

    The Late-Breaking Demo accepted at ISMIR 2026 tests whether a 103k parameter student, distilled from music2latent, keeps the codec's similarity ordering, and what then limits real time operation on the Elk Stomp.

    Insights

    • Spearman ρ .829 for the student, against .797 for a ridge map and .575 for MFCCs, over 48 query and corpus pairs
    • Its picks rank near the teacher's own second seed (percentile .945 against .981)
    • On the pedal the student takes about 1% of each 93 ms analysis; only the search grows with the corpus, and an 8 bit table speeds it 2.6 to 3.1 times
  3. Release Timed on the board, then made faster

    Per stage timers on the pedal's worker, an 8 bit matching table, and a front end that no longer divides 64 bit integers.

    Customer value. Bigger sound worlds fit the 93 ms analysis period, and the paper's on device numbers come straight from these runs.

    • Front end: 16.7 → 4.0 ms per hop
    • 8 bit table: search 2.6 to 3.1 times faster, ordering unchanged (ρ 1.000)
    • 12,018 grain hop: 30.2 → 17.6 ms
    • Golden parity held at 200/200; offline bounces byte identical
  4. Release Five sound worlds, switched live from the panel

    The rotary picks between five corpora on the pedal, each loaded on a helper thread and swapped without touching the audio thread.

    Customer value. One pedal, several instruments: a performer changes sound world midset, with no laptop.

    • Live switching verified four times on the board, zero drops
    • Matching spread across hops: board p99 340 → 95 ms, picks bit identical
    • Two playing sessions on the board (10 and 15 minutes), clean
    • World select on the rotary, output level on a pot
  5. Release The instrument runs on the pedal

    Cross built for the Elk Stomp's 32 bit ARM cores and loaded under SUSHI without the vendor SDK; the answer bar now comes from the pedal's own output jacks.

    Customer value. No laptop on stage: the instrument the challenge asked for, played from the pedal by feel.

    • Docker cross build: Ubuntu 24.04 armhf, GCC 13, JUCE 8
    • Checksummed file transfer over the serial console, about 8 KB/s
    • On the board: worker p50 93.3 ms, p99 94.9 ms, zero overruns, stalls or drops
    • All five pot mappings and hold verified by hand

    Risks

    • Per stage costs on the board still unmeasured
  6. Release The public talk at MUTEK, told straight

    The team's Winners Showcase presentation for the Elk Audio Challenge (Grands Ballets, August 28): THIRI as a product, and the honest account of a desktop plugin meeting a hard real time pedal.

    Customer value. Musicians hear "play one horn, hear five; it fits in a stompbox"; engineers get the measured path from 13× over deadline to five voices at 0.59× with headroom.

    • THIRI as a family: standalone, VST, and the live demo at demo.thiri.ai
    • The porting lesson, on the slide: know the specs before you rebuild
    • Future directions: ML on device via distillation; no GPU on the fly model building

    Risks

    • Sponsor and festival marks are attribution, not endorsement
  7. Release The instrument runs under Elk's engine

    End to end on desktop, under the same engine the pedal runs: live input → student embedding → grain match over a 12,018 grain corpus → the answer, one bar late.

    Customer value. The whole instrument is testable and audible before a single cross compile; configs and knob mappings transfer to hardware unchanged.

    • Corpus pipeline: recordings → one pack file (18.6 min in 18 s, 84× real time)
    • Golden parity: 200/200 identical grain selections vs the Python reference
    • Async worker under the host: p99 cadence <99 ms, zero drops; audio thread 0.09%
    • Ten knob mapped parameters; bar clock on the plugin's own sample counter

    Risks

    • A7 budgets are projected until the board benchmarks run
  8. Research · Experiment Distillation judged by regret, not agreement

    Can a 103k parameter student stand in for a 58M parameter codec at the instrument's front end? Four conditions, four corpora, held out queries, and a metric correction.

    Insights

    • The oracle agrees with itself on only 2.6% of picks at live settings; regret is the honest lens, not agreement
    • The student lands in the oracle's top 4–7% (ρ .84); a raw MFCC baseline fails at 15–18%
    • A recurrent variant tied the simple MLP everywhere; the stateless model shipped
  9. Research · Synthesis What actually fits on a pedal with no NPU

    An eight chapter, cited knowledge base of the Elk Stomp platform (hardware, OS, host internals, and the real budget for machine learning), verified claim by claim.

    Insights

    • The budget gap is categorical: ~2,000 MACs per sample at audio rate, millions per frame at 10.7 Hz
    • No PyTorch runtime exists for the board's 32 bit architecture; inference compiles into the plugin
    • Two latent host defects found by reading its source, both now design constraints
  10. Incident · Sev2 Resolved Loaded fine, averaged 6.8%, peaked at 1302%

    The team's plugin cross compiled, loaded on the Elk Stomp with all parameters live, then overloaded the 1.333 ms audio deadline the moment its DSP engaged.

    Impact. No on device demo at the event. The board's telemetry: average 6.8% of budget, peak 1302%, roughly 13× over.

    Detection. Audible glitching under load; confirmed postevent from the engine's per processor timing statistics recovered off the board.

    Root cause. Frame based DSP built around 512 sample hops concentrates work into one callback in eight. The deadline is per callback peak, not average.

    Fix. The team replaced the spectral shifter with time domain PSOLA (~50× cheaper per voice): five voices at 0.59× of the deadline. This instrument made the rule architecture: analysis on a 10.7 Hz worker, the audio thread only mixes, an allocation guard in debug.

    Followups

    • Fold "peak, not average" into the realtime review checklist (done)
    • Verify no involuntary mode switches on the pedal (board session) (open)

    Blameless note. Desktop plugin conventions assume block sizes this platform does not offer. The measurement became a design rule, a review gate, and a slide in the public talk.

  11. Research · Experiment Prehackathon: a folder of recordings becomes an instrument

    The week before the hackathon, the latent granular idea was built and gated on a laptop: encode a corpus, match live playing grain by grain in the latent space, answer a bar late.

    Insights

    • The latency became the identity: one block ahead reads as an answer, not a lag
    • Antihub and stickiness penalties matter as much as similarity (worst corpus, 88% of picks from five grains)
    • Five shaping controls were enough; presets mattered more than a sixth
12 · Decisions

Decision records

Accepted Play real grains instead of running a decoder

Context. The laptop instrument decodes latents back to audio through the codec: the most expensive stage, and the one a pedal with no NPU most clearly refuses.

Decision. Delete the decoder. Match in the learned space; play back int16 corpus audio with equal power crossfades from precomputed margins.

Why. The mosaic character (real grains, intelligently chosen) is the half audiences respond to, and it survives intact. Intelligence in the choosing, not the manufacturing.

Consequences. The audio thread only mixes and windows. Morph is a scoped out feature with a named path back, not a silent omission.

Accepted Move the decisions, not the model

Context. The instrument depends on a 58M parameter codec that is slower than real time even on a laptop CPU. The target is two 800 MHz cores, no NPU, and no PyTorch runtime exists for its 32 bit architecture at all.

Decision. Restructure. The codec embeds the corpus once, offline; a distilled 103k parameter student embeds live input at 10.7 Hz on a worker thread. The audio thread never runs a network.

Why. At audio rate the board affords ~2,000 multiply accumulates per sample; at the decision rate, millions per frame. Same silicon, three orders of magnitude apart.

Consequences. Student inference costs 1–2% of one core; the audio thread measures 0.09%. The quality question moved somewhere measurable (the regret study), and the recipe generalizes to any latency critical system with intermittent decisions.