Verifying Deep Keyword Spotting Detection with Acoustic Word Embeddings - Citegraph

Paper Info

Title
Verifying Deep Keyword Spotting Detection with Acoustic Word Embeddings

Abstract
In this paper, in order to improve keyword spotting (KWS) performance in a live broadcast scenario, we propose to use a template matching method based on acoustic word embeddings (AWE) as the second stage to verify the detection from the Deep KWS system. AWEs are obtained via a deep bidirectional long short-term memory (BLSTM) network trained using limited positive and negative keyword candidates, which aims to encode variable-length keyword candidates into fixed-dimensional vectors with reasonable discriminative ability. Learning AWEs takes a combination of three specifically-designed losses: the triplet and reversed triplet losses try to keep same keyword candidates closer and different keyword candidates farther, while the hinge loss is to set a fixed threshold to distinguish all positive and negative keyword candidates. During keyword verification, calibration scores are used to reduce the bias between different templates for different keyword candidates. Experiments show that adding AWE-based keyword verification to Deep KWS achieves 5.6% relative accuracy improvement; the hinge loss brings additional 5.5% relative gain and the final accuracy climbs to 0.775 by using calibration scores.

Year	DOI	Venue
2019	10.1109/ASRU46091.2019.9003781	2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Keywords	DocType	ISBN
Query-by-example,keyword spotting,acoustic word embeddings,hinge loss,calibration scores	Conference	978-1-7281-0307-5
Citations	PageRank	References
0	0.34	0
Authors
4

Authors (4 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Yougen Yuan	1	8	2.78
Zhiqiang Lv	2	26	11.28
Shen Huang	3	64	14.51
Lei Xie	4	425	64.87

1