Projects
Generative audio

SoloMuse

Listens to a backing track and generates a matching, steerable solo as real audio.

Generative audioTransformersPyTorchEnCodec

What it is

SoloMuse listens to a backing track and generates a solo that fits it, as actual audio. Unlike a black-box model, it plans the performance first (when to play, how dense, how tense, how loud), so the output can be steered.

How it works

  1. Situation layer. Reads the backing track (tempo, loudness, harmony/chroma, spectral features) into a compact 32-number summary.
  2. Intent planner. A causal Transformer plans the musical intent 10 times per second: play or rest, note density, register, tension, dynamics, onsets and phrasing. You can override it, e.g. force more tension.
  3. Renderer. An autoregressive multi-codebook Transformer language model generates Meta EnCodec audio tokens (4 codebooks × 1,024 tokens at 75 Hz), which are decoded into a waveform.
SoloMuse architecture diagram
System architecture: situation layer → intent planner → EnCodec token renderer.

Data pipeline

Engineering highlights

127automated tests
3model layers
A100cloud GPUs on RunPod
250 msstreaming chunks

Built with

PyTorchMeta EnCodecLibrosapyloudnormDemucsDockerRunPod
View on GitHub
All projects Ask Gizmo