add preprocess requirements

This commit is contained in:
钱家乐
2026-02-06 22:12:13 +08:00
parent e0cf305dee
commit 20bceb3985
2 changed files with 65 additions and 2 deletions
+32 -2
View File
@@ -20,9 +20,35 @@ The toolkit includes the following core modules:
Supports customizable creation and editing of MIDI scores integrated with lyrics.
## 📁 Data Preparation
## 🔧 Python Environment
To ensure the data processing pipeline runs correctly, please verify that all required checkpoints are correctly placed in `pretrained_models/SoulX-Singer-Preprocess`
Before running the pipeline, set up the Python environment as follows:
1. **Install Conda** (if not already installed): https://docs.conda.io/en/latest/miniconda.html
2. **Activate or create a conda environment** (recommended Python 3.10):
- If you already have the `soulxsinger` environment:
```bash
conda activate soulxsinger
```
- Otherwise, create it first:
```bash
conda create -n soulxsinger -y python=3.10
conda activate soulxsinger
```
3. **Install dependencies** from the `preprocess` directory:
```bash
cd preprocess
pip install -r requirements.txt
```
## 📁 Data Preparation
Before running the pipeline, prepare the following inputs:
@@ -61,6 +87,10 @@ The script will automatically execute the following steps:
After the pipeline completes, you will obtain **SoulX-Singer–style metadata** that can be directly used for Singing Voice Synthesis (SVS).
**Output paths:**
- The final metadata (**JSON file**) is written **in the same directory as your input audio**, with the **same filename** (e.g. `audio.mp3` → `audio.json`)
- All **intermediate results** (separated vocal and accompaniment, F0, VAD outputs, etc.) are also saved under the configured **`save_dir`**.
⚠️ **Important Note**
Transcription errors—especially in **lyrics** and **note annotations**—can significantly affect the final SVS quality. We **strongly recommend manually reviewing and correcting** the generated metadata before inference.
+33
View File
@@ -0,0 +1,33 @@
beartype==0.22.9
einops==0.8.2
funasr==1.3.0
g2p_en==2.1.0
g2pM==0.1.2.5
librosa==0.11.0
loralib==0.1.2
matplotlib==3.10.8
mido==1.3.3
ml_collections==1.1.0
nemo_toolkit==2.6.1
nltk==3.9.2
numba==0.63.1
numpy==2.2.6
omegaconf==2.3.0
packaging==24.2
praat-parselmouth==0.4.7
pretty_midi==0.2.11
pyloudnorm==0.2.0
pyworld==0.3.5
rotary_embedding_torch==0.8.9
sageattention==1.0.6
scikit_learn==1.7.2
scipy==1.15.3
six==1.17.0
scikit_image==0.25.2
soundfile==0.13.1
ToJyutping==3.2.0
torch==2.10.0
torchaudio==2.10.0
tqdm==4.67.1
wandb==0.24.2
webrtcvad==2.0.10