add preprocess requirements
This commit is contained in:
+32
-2
@@ -20,9 +20,35 @@ The toolkit includes the following core modules:
|
||||
Supports customizable creation and editing of MIDI scores integrated with lyrics.
|
||||
|
||||
|
||||
## 📁 Data Preparation
|
||||
## 🔧 Python Environment
|
||||
|
||||
To ensure the data processing pipeline runs correctly, please verify that all required checkpoints are correctly placed in `pretrained_models/SoulX-Singer-Preprocess`
|
||||
Before running the pipeline, set up the Python environment as follows:
|
||||
|
||||
1. **Install Conda** (if not already installed): https://docs.conda.io/en/latest/miniconda.html
|
||||
|
||||
2. **Activate or create a conda environment** (recommended Python 3.10):
|
||||
|
||||
- If you already have the `soulxsinger` environment:
|
||||
|
||||
```bash
|
||||
conda activate soulxsinger
|
||||
```
|
||||
|
||||
- Otherwise, create it first:
|
||||
|
||||
```bash
|
||||
conda create -n soulxsinger -y python=3.10
|
||||
conda activate soulxsinger
|
||||
```
|
||||
|
||||
3. **Install dependencies** from the `preprocess` directory:
|
||||
|
||||
```bash
|
||||
cd preprocess
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
## 📁 Data Preparation
|
||||
|
||||
Before running the pipeline, prepare the following inputs:
|
||||
|
||||
@@ -61,6 +87,10 @@ The script will automatically execute the following steps:
|
||||
|
||||
After the pipeline completes, you will obtain **SoulX-Singer–style metadata** that can be directly used for Singing Voice Synthesis (SVS).
|
||||
|
||||
**Output paths:**
|
||||
- The final metadata (**JSON file**) is written **in the same directory as your input audio**, with the **same filename** (e.g. `audio.mp3` → `audio.json`)
|
||||
- All **intermediate results** (separated vocal and accompaniment, F0, VAD outputs, etc.) are also saved under the configured **`save_dir`**.
|
||||
|
||||
⚠️ **Important Note**
|
||||
|
||||
Transcription errors—especially in **lyrics** and **note annotations**—can significantly affect the final SVS quality. We **strongly recommend manually reviewing and correcting** the generated metadata before inference.
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
beartype==0.22.9
|
||||
einops==0.8.2
|
||||
funasr==1.3.0
|
||||
g2p_en==2.1.0
|
||||
g2pM==0.1.2.5
|
||||
librosa==0.11.0
|
||||
loralib==0.1.2
|
||||
matplotlib==3.10.8
|
||||
mido==1.3.3
|
||||
ml_collections==1.1.0
|
||||
nemo_toolkit==2.6.1
|
||||
nltk==3.9.2
|
||||
numba==0.63.1
|
||||
numpy==2.2.6
|
||||
omegaconf==2.3.0
|
||||
packaging==24.2
|
||||
praat-parselmouth==0.4.7
|
||||
pretty_midi==0.2.11
|
||||
pyloudnorm==0.2.0
|
||||
pyworld==0.3.5
|
||||
rotary_embedding_torch==0.8.9
|
||||
sageattention==1.0.6
|
||||
scikit_learn==1.7.2
|
||||
scipy==1.15.3
|
||||
six==1.17.0
|
||||
scikit_image==0.25.2
|
||||
soundfile==0.13.1
|
||||
ToJyutping==3.2.0
|
||||
torch==2.10.0
|
||||
torchaudio==2.10.0
|
||||
tqdm==4.67.1
|
||||
wandb==0.24.2
|
||||
webrtcvad==2.0.10
|
||||
Reference in New Issue
Block a user