Text-Free Prosody-Aware Generative Spoken Language Modeling - Citegraph

Paper Info

Title
Text-Free Prosody-Aware Generative Spoken Language Modeling

Abstract
Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generative Spoken Language Modeling (GSLM) (Lakhotia et al., 2021) is the only prior work addressing the generative aspects of speech pre-training, which replaces text with discovered phone-like units for language modeling and shows the ability to generate meaningful novel sentences. Unfortunately, despite eliminating the need of text, the units used in GSLM discard most of the prosodic information. Hence, GSLM fails to leverage prosody for better comprehension, and does not generate expressive speech. In this work, we present a prosody-aware generative spoken language model (pGSLM). It is composed of a multi-stream transformer language model (MS-TLM) of speech, represented as discovered unit and prosodic feature streams, and an adapted HiFi-GAN model converting MS-TLM outputs to waveforms. We devise a series of metrics for prosody modeling and generation, and re-use metrics from GSLM for content modeling. Experimental results show that the pGSLM can utilize prosody to improve both prosody and content modeling, and also generate natural, meaningful, and coherent speech given a spoken prompt.(1)

Year	DOI	Venue
2022	10.18653/v1/2022.acl-long.593	PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS)
DocType	Volume	Citations
Conference	Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)	2
PageRank	References	Authors
0.35	0	11

Authors (11 rows)

Cited by (2 rows)

References (0 rows)

Name	Order	Citations	PageRank
Eugene Kharitonov	1	67	6.63
Ann B Lee	2	602	56.97
Adam Polyak	3	50	6.09
Yossi Adi	4	87	9.18
Jade Copet	5	4	0.70
Kushal Lakhotia	6	4	0.70
Tu Anh T. Nguyen	7	56	9.27
Morgane Rivière	8	8	2.54
Abdel-rahman Mohamed	9	3772	266.13
Emmanuel Dupoux	10	238	37.33
Wei-Ning Hsu	11	115	13.93

1