Mutual-Optimization Towards Generative Adversarial Networks For Robust Speech Recognition - Citegraph

Paper Info

Title
Mutual-Optimization Towards Generative Adversarial Networks For Robust Speech Recognition

Abstract
In the context of Automatic Speech Recognition (ASR), improving the noise robustness remains an intractable task. Speech enhancement, combined with Generative Adversarial Networks (GAN), such as SEGAN, has effective performance in denoising raw waveform speech signals. Instead of waveforms, using Mel filterbank spectra in GAN is proposed, which has better performance in the task of ASR. However, these techniques will still miss useful information when GAN is used in them. In this paper, we investigate to protect the useful information in GAN, and propose a novel model, called Discriminator Generator Classifier-GAN (DGC-GAN). While normal GAN combining just two networks will lead the model to denoising rather than recognition, DGC-GAN has another network called classifier, which is an ASR system that will tune GAN to be recognized easier. By adding a classifier into previous GAN to get DGC-GAN, we achieve 29.1% Phone Error Rate (PER) relative improvement in a tiny dataset and 47.4% PER relative improvement in a large dataset.

Year	DOI	Venue
2018	10.1109/ICPR.2018.8546090	2018 24TH INTERNATIONAL CONFERENCE ON PATTERN RECOGNITION (ICPR)
Keywords	Field	DocType
automatic speech recognition, speech enhancement, Mel filterbank spectra, generative adversarial networks	Noise reduction,Speech enhancement,Discriminator,Noise measurement,Pattern recognition,Computer science,Filter bank,Word error rate,Speech recognition,Robustness (computer science),Artificial intelligence,Classifier (linguistics)	Conference
ISSN	Citations	PageRank
1051-4651	0	0.34
References	Authors
0	5

Authors (5 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Ke Ding	1	0	1.69
Ne Luo	2	0	0.34
Yanyan Xu	3	15	7.28
Dengfeng Ke	4	12	6.51
Kaile Su	5	668	60.11

1