Title | ||
---|---|---|
Quantifying the value of pronunciation lexicons for keyword search in lowresource languages |
Abstract | ||
---|---|---|
This paper quantifies the value of pronunciation lexicons in large vocabulary continuous speech recognition (LVCSR) systems that support keyword search (KWS) in low resource languages. State-of-the-art LVCSR and KWS systems are developed for conversational telephone speech in Tagalog, and the baseline lexicon is augmented via three different grapheme-to-phoneme models that yield increasing coverage of a large Tagalog word-list. It is demonstrated that while the increased lexical coverage - or reduced out-of-vocabulary (OOV) rate - leads to only modest (ca 1%-4%) improvements in word error rate, the concomitant improvements in actual term weighted value are as much as 60%. It is also shown that incorporating the augmented lexicons into the LVCSR system before indexing speech is superior to using them post facto, e.g., for approximate phonetic matching of OOV keywords in pre-indexed lattices. These results underscore the disproportionate importance of automatic lexicon augmentation for KWS in morphologically rich languages, and advocate for using them early in the LVCSR stage. |
Year | DOI | Venue |
---|---|---|
2013 | 10.1109/ICASSP.2013.6639336 | ICASSP |
Keywords | Field | DocType |
keyword search,information retrieval,lexical coverage,speech recognition,pronunciation lexicons value,large vocabulary continuous speech recognition,word error rate,indexing speech,out-of-vocabulary rate,indexing,tagalog word-list,kws systems,approximate phonetic matching,concomitant improvements,morphology,lvcsr systems,conversational telephone speech,preindexed lattices,low resource languages,speech synthesis,grapheme-to-phoneme models,baseline lexicon,speech,acoustics,lattices,hidden markov models | Tagalog,Speech corpus,Pronunciation,Computer science,Keyword search,Word error rate,Search engine indexing,Speech recognition,Lexicon,Natural language processing,Artificial intelligence,Vocabulary | Conference |
ISSN | Citations | PageRank |
1520-6149 | 16 | 1.19 |
References | Authors | |
11 | 6 |
Name | Order | Citations | PageRank |
---|---|---|---|
Guoguo Chen | 1 | 428 | 19.89 |
Sanjeev Khudanpur | 2 | 2155 | 202.00 |
Daniel Povey | 3 | 2442 | 231.75 |
Jan Trmal | 4 | 235 | 20.91 |
David Yarowsky | 5 | 3986 | 618.81 |
Oguz Yilmaz | 6 | 52 | 2.61 |