Title
Exploiting deep neural networks for detection-based speech recognition.
Abstract
In recent years deep neural networks (DNNs) - multilayer perceptrons (MLPs) with many hidden layers - have been successfully applied to several speech tasks, i.e., phoneme recognition, out of vocabulary word detection, confidence measure, etc. In this paper, we show that DNNs can be used to boost the classification accuracy of basic speech units, such as phonetic attributes (phonological features) and phonemes. This boosting leads to higher flexibility and has the potential to integrate both top-down and bottom-up knowledge into the Automatic Speech Attribute Transcription (ASAT) framework. ASAT is a new family of lattice-based speech recognition systems grounded on accurate detection of speech attributes. In this paper we compare DNNs and shallow MLPs within the ASAT framework to classify phonetic attributes and phonemes. Several DNN architectures ranging from five to seven hidden layers and up to 2048 hidden units per hidden layer will be presented and evaluated. Experimental evidence on the speaker-independent Wall Street Journal corpus clearly demonstrates that DNNs can achieve significant improvements over the shallow MLPs with a single hidden layer, producing greater than 90% frame-level attribute estimation accuracies for all 21 phonetic features tested. Similar improvement is also observed on the phoneme classification task with excellent frame-level accuracy of 86.6% by using DNNs. This improved phoneme prediction accuracy, when integrated into a standard large vocabulary continuous speech recognition (LVCSR) system through a word lattice rescoring framework, results in improved word recognition accuracy, which is better than previously reported word lattice rescoring results.
Year
DOI
Venue
2013
10.1016/j.neucom.2012.11.008
Neurocomputing
Keywords
Field
DocType
deep neural network,basic speech unit,large vocabulary continuous speech,shallow mlps,hidden layer,detection-based speech recognition,single hidden layer,phonetic attribute,hidden unit,lattice-based speech recognition system,speech attribute,word lattice,speech recognition
Automatic speech,Pattern recognition,Computer science,Word recognition,Speech recognition,Ranging,Artificial intelligence,Boosting (machine learning),Perceptron,Vocabulary,Phoneme recognition,Deep neural networks
Journal
Volume
ISSN
Citations 
106
0925-2312
33
PageRank 
References 
Authors
0.93
49
4
Name
Order
Citations
PageRank
Sabato Marco Siniscalchi131030.21
Dong Yu26264475.73
Deng, Li39691728.14
Chin-Hui Lee46101852.71