Self-directed
Music Splitter

DSP from scratch
serving
Project overview
From-scratch music source separation
Upload a song, get back four isolated stems, vocals, drums, bass, and other, each playable and downloadable. The entire signal-processing and ML stack is hand-implemented in PyTorch, no separation library.
Date
2026
Field
AI / ML
Stack
Python, PyTorch, FastAPI, torchaudio, NumPy
4 isolated stems
STFT reconstruction to ~1e-6
41 CPU tests
Music Splitter
Context
A waveform-domain Demucs U-Net with a generalized Wiener post-filter, STFT/iSTFT, complex masking, and SI-SDR.

Most separation projects glue together a library and a checkpoint.
The hard part
I wanted to derive the math myself, the STFT as a single matmul, the Wiener filter, the loss, to actually understand source separation end to end.


What it took
- Implemented the STFT/iSTFT as strided framing plus a precomputed complex DFT basis matmul, satisfying NOLA/COLA for perfect reconstruction to ~1e-6 and energy conservation to 1e-4, validated against torch.stft.
- Built a waveform-domain Demucs U-Net (~34M params) with a BiLSTM bottleneck and U-Net skips, trained against a negative per-stem SI-SDR objective with an optional permutation-invariant wrapper.
- Added a generalized Wiener filter forming a partition of unity that recovers oracle sources at 28–36 dB SI-SDR, plus 8-bit quantization-aware training with a straight-through estimator for ~50% size reduction.





