/work/music-splitter
AmanShah
Brief
amanashishshah@gmail.com

© 2026 Aman Shah

Recruiter mode

Self-directed

Work/

Music Splitter

Music Splitter, at a glance

DSP from scratch

serving

Project overview

From-scratch music source separation

Upload a song, get back four isolated stems, vocals, drums, bass, and other, each playable and downloadable. The entire signal-processing and ML stack is hand-implemented in PyTorch, no separation library.

Role

  • ML engineer
  • DSP and model implementation
  • serving

Elsewhere

Date

2026

Field

AI / ML

Stack

Python, PyTorch, FastAPI, torchaudio, NumPy

4 isolated stems

STFT reconstruction to ~1e-6

41 CPU tests

Music Splitter

Context

A waveform-domain Demucs U-Net with a generalized Wiener post-filter, STFT/iSTFT, complex masking, and SI-SDR.

Music Splitter, image 2

Most separation projects glue together a library and a checkpoint.

The hard part

I wanted to derive the math myself, the STFT as a single matmul, the Wiener filter, the loss, to actually understand source separation end to end.

Music Splitter, image 4
Music Splitter, image 4

What it took

  • Implemented the STFT/iSTFT as strided framing plus a precomputed complex DFT basis matmul, satisfying NOLA/COLA for perfect reconstruction to ~1e-6 and energy conservation to 1e-4, validated against torch.stft.
  • Built a waveform-domain Demucs U-Net (~34M params) with a BiLSTM bottleneck and U-Net skips, trained against a negative per-stem SI-SDR objective with an optional permutation-invariant wrapper.
  • Added a generalized Wiener filter forming a partition of unity that recovers oracle sources at 28–36 dB SI-SDR, plus 8-bit quantization-aware training with a straight-through estimator for ~50% size reduction.
Music Splitter, image 6

Outcome

A FastAPI app that separates arbitrary-length tracks in cross-faded overlapping chunks with bounded memory, switchable between a from-scratch and a pretrained Demucs engine and an optional Wiener post-filter, behind an editorial 'the math as content' UI.