Python MIT

vocal-remover

Vocal Remover using Deep Neural Networks

T

tsurumeso

Dernière activité 10 sept. 2026
tsurumeso/vocal-remover

1,8 k

étoiles

256

forks

69

issues ouvertes

audiodeep-learningpytorchsegmentationspectrogramvocal-removervocal-separation

Ce README est souvent en anglais.

vocal-remover

Release Release

This is a deep-learning-based tool for extracting the instrumental track from your songs.

Installation

Getting vocal-remover

Download the latest version from here.

Install PyTorch

See: GET STARTED

Install the other packages

cd vocal-remover
pip install -r requirements.txt

Usage

The following command separates the input into instrumental and vocal tracks. They are saved as *_Instruments.wav and *_Vocals.wav.

Run on CPU

python inference.py --input path/to/an/audio/file

Run on GPU

python inference.py --input path/to/an/audio/file --gpu 0

Advanced options

The --tta option performs Test-Time Augmentation to improve separation quality.

python inference.py --input path/to/an/audio/file --tta --gpu 0

The --postprocess option masks the instrumental track based on the vocal volume to improve separation quality.

Warning

This is an experimental feature. If you encounter any problems with this option, please disable it.

python inference.py --input path/to/an/audio/file --postprocess --gpu 0

Train your own model

Place your dataset

path/to/dataset/
  +- instruments/
  |    +- 01_foo_inst.wav
  |    +- 02_bar_inst.mp3
  |    +- ...
  +- mixtures/
       +- 01_foo_mix.wav
       +- 02_bar_mix.mp3
       +- ...

Train a model

python train.py --dataset path/to/dataset --mixup_rate 0.5 --reduction_rate 0.5 --gpu 0

References

Projets similaires

The PyTorch-based audio source separation toolkit for researchers

Pythonaudio-separationdeep-learningpretrained-models
Aasteroid-team
2,6 k étoiles452

Data manipulation and transformation for audio signal processing, powered by PyTorch

Pythonaudioaudio-processingio
Ppytorch
3 k étoiles800

A PyTorch-based Speech Toolkit

Pythonasraudioaudio-processing
Sspeechbrain
11,8 k étoiles1,7 k