ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders - Citegraph

Paper Info

Title
ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders

Abstract
This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders. Different from the conventional SVS models, the proposed ByteSing employs Tacotron-like encoder-decoder structures as the acoustic models, in which the CBHG models and recurrent neural networks (RNNs) are explored as encoders and decoders respectively. Meanwhile an auxiliary phoneme duration prediction model is utilized to expand the input sequence, which can enhance the model controllable capacity, model stability and tempo prediction accuracy. WaveRNN vocoders are also adopted as neural vocoders to further improve the voice quality of synthesized songs. Both objective and subjective experimental results prove that the SVS method proposed in this paper can produce quite natural, expressive and high-fidelity songs by improving the pitch and spectrogram prediction accuracy and the models using attention mechanism can achieve best performance.

Year	DOI	Venue
2021	10.1109/ISCSLP49672.2021.9362104	2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)
Keywords	DocType	ISBN
ByteSing,Singing voice synthesis,Tacotron,WaveRNN,Duration allocated	Conference	978-1-7281-6995-8
Citations	PageRank	References
0	0.34	0
Authors
9

Authors (9 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Gu Yu	1	0	0.34
Yin Xiang	2	0	1.01
Rao Yonghui	3	0	0.34
Wan Yuan	4	0	0.34
Tang Benlai	5	0	0.34
Zhang Yang	6	0	0.34
Jitong Chen	7	199	10.01
Yu-Xuan Wang	8	650	32.68
Ma Zejun	9	0	0.34

1