Exploiting deep neural networks for detection-based speech recognition. - Citegraph

Paper Info

Title
Exploiting deep neural networks for detection-based speech recognition.

Abstract
In recent years deep neural networks (DNNs) - multilayer perceptrons (MLPs) with many hidden layers - have been successfully applied to several speech tasks, i.e., phoneme recognition, out of vocabulary word detection, confidence measure, etc. In this paper, we show that DNNs can be used to boost the classification accuracy of basic speech units, such as phonetic attributes (phonological features) and phonemes. This boosting leads to higher flexibility and has the potential to integrate both top-down and bottom-up knowledge into the Automatic Speech Attribute Transcription (ASAT) framework. ASAT is a new family of lattice-based speech recognition systems grounded on accurate detection of speech attributes. In this paper we compare DNNs and shallow MLPs within the ASAT framework to classify phonetic attributes and phonemes. Several DNN architectures ranging from five to seven hidden layers and up to 2048 hidden units per hidden layer will be presented and evaluated. Experimental evidence on the speaker-independent Wall Street Journal corpus clearly demonstrates that DNNs can achieve significant improvements over the shallow MLPs with a single hidden layer, producing greater than 90% frame-level attribute estimation accuracies for all 21 phonetic features tested. Similar improvement is also observed on the phoneme classification task with excellent frame-level accuracy of 86.6% by using DNNs. This improved phoneme prediction accuracy, when integrated into a standard large vocabulary continuous speech recognition (LVCSR) system through a word lattice rescoring framework, results in improved word recognition accuracy, which is better than previously reported word lattice rescoring results.

Year	DOI	Venue
2013	10.1016/j.neucom.2012.11.008	Neurocomputing
Keywords	Field	DocType
deep neural network,basic speech unit,large vocabulary continuous speech,shallow mlps,hidden layer,detection-based speech recognition,single hidden layer,phonetic attribute,hidden unit,lattice-based speech recognition system,speech attribute,word lattice,speech recognition	Automatic speech,Pattern recognition,Computer science,Word recognition,Speech recognition,Ranging,Artificial intelligence,Boosting (machine learning),Perceptron,Vocabulary,Phoneme recognition,Deep neural networks	Journal
Volume	ISSN	Citations
106	0925-2312	33
PageRank	References	Authors
0.93	49	4

Authors (4 rows)

Cited by (33 rows)

References (49 rows)

Name	Order	Citations	PageRank
Sabato Marco Siniscalchi	1	310	30.21
Dong Yu	2	6264	475.73
Deng, Li	3	9691	728.14
Chin-Hui Lee	4	6101	852.71

1