Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions

Rong Gong; Xavier Serra

Note: This bibliographic page is archived and will no longer be updated. For an up-to-date list of publications from the Music Technology Group see the Publications list .

Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions

Title	Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions
Publication Type	Conference Paper
Year of Publication	2018
Conference Name	Interspeech 2018
Authors	Gong, R. , & Serra X.
Conference Start Date	02/09/2018
Conference Location	Hyderabad, India
Abstract	In this paper, we tackle the singing voice phoneme segmentation problem in the singing training scenario by using language-independent information -- onset and prior coarse duration. We propose a two-step method. In the first step, we jointly calculate the syllable and phoneme onset detection functions (ODFs) using a convolutional neural network (CNN). In the second step, the syllable and phoneme boundaries and labels are inferred hierarchically by using a duration-informed hidden Markov model (HMM). To achieve the inference, we incorporate the a priori duration model as the transition probabilities and the ODFs as the emission probabilities into the HMM. The proposed method is designed in a language-independent way such that no phoneme class labels are used. For the model training and algorithm evaluation, we collect a new jingju (also known as Beijing or Peking opera) solo singing voice dataset and manually annotate the boundaries and labels at phrase, syllable and phoneme levels. The dataset is publicly available. The proposed method is compared with a baseline method based on hidden semi-Markov model (HSMM) forced alignment. The evaluation results show that the proposed method outperforms the baseline by a large margin regarding both segmentation and onset detection tasks.
preprint/postprint document	https://arxiv.org/abs/1806.01665

Additional material:

Jingju solo singing voice audio and annotation dataset: https://doi.org/10.5281/zenodo.1185123

Code and supplementary materials: https://github.com/ronggong/interspeech2018_submission01

Jupyter notebook demo running in Google Colab: https://goo.gl/BzajRy