Acoustical and perceptual study of voice disguise by age modification in speaker verification. - Citegraph

Paper Info

Title
Acoustical and perceptual study of voice disguise by age modification in speaker verification.

Abstract
The task of speaker recognition is feasible when the speakers are co-operative or wish to be recognized. While modem automatic speaker verification (ASV) systems and some listeners are good at recognizing speakers from modal, unmodified speech, the task becomes notoriously difficult in situations of deliberate voice disguise when the speaker aims at masking his or her identity. We approach voice disguise from the perspective of acoustical and perceptual analysis using a self-collected corpus of 60 native Finnish speakers (31 female, 29 male) producing utterances in normal, intended young and intended old voice modes. The normal voices form a starting point and we are interested in studying how the two disguise modes impact the acoustical parameters and perceptual speaker similarity judgments. First, we study the effect of disguise as a relative change in fundamental frequency (F0) and formant frequencies (F1 to P4) from modal to disguised utterances. Next, we investigate whether or not speaker comparisons that are deemed easy or difficult by a modern ASV system have a similar difficulty level for the human listeners. Further, we study affecting factors from listener-related self-reported information that may explain a particular listener's success or failure in speaker similarity assessment. Our acoustic analysis reveals a systematic increase in relative change in mean FO for the intended young voices while for the intended old voices, the relative change is less prominent in most cases. Concerning the formants F1 through F4, 29% (for male) and 30% (for female) of the utterances did not exhibit a significant change in any formant value, while the remaining 70% of utterances had significant changes in at least one formant. Our listening panel consists of 70 listeners, 32 native and 38 non-native, who listened to 24 utterance pairs selected using rankings produced by an ASV system. The results indicate that speaker pairs categorized as easy by our ASV system were also easy for the average listener. Similarly, the listeners made more errors in the difficult trials. The listening results indicate that target (same speaker) trials were more difficult for the non-native group, while the performance for the non-target pairs was similar for both native and non-native groups.

Year	DOI	Venue
2017	10.1016/j.specom.2017.10.002	Speech Communication
Keywords	Field	DocType
Voice disguise,Voice modification,Speaker verification,Acoustical analysis,Fundamental frequency,Formant frequencies,Perceptual evaluation	Speaker verification,Computer science,Utterance,Active listening,Speech recognition,Speaker recognition,Speaker diarisation,Formant,Perception,Modal	Journal
Volume	ISSN	Citations
95	0167-6393	4
PageRank	References	Authors
0.43	17	4

Authors (4 rows)

Cited by (4 rows)

References (17 rows)

Name	Order	Citations	PageRank
Rosa González Hautamäki	1	30	3.87
Md. Sahidullah	2	326	24.99
Ville Hautamäki	3	385	33.51
Tomi Kinnunen	4	1323	86.67

1