Title
Comparing intermittency and network measurements of words and their dependency on authorship
Abstract
Many features of texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper, we quantify how topological properties of word co-occurrence networks and intermittency (or burstiness) in word distribution depend on the style of authors. Our database contains 40 books by eight authors who lived in the nineteenth and twentieth centuries, for which the following network measurements were obtained: the clustering coefficient, average shortest path lengths and betweenness. We found that the two factors with stronger dependence on authors were skewness in the distribution of word intermittency and the average shortest paths. Other factors such as betweenness and Zipf's law exponent show only weak dependence on authorship. Also assessed was the contribution from each measurement to authorship recognition using three machine learning methods. The best performance was about 65% accuracy upon combining complex networks and intermittency features with the nearest-neighbor algorithm of automatic authorship. From a detailed analysis of the interdependence of the various metrics, it is concluded that the methods used here are complementary for providing short- and long-scale perspectives on texts, which are useful for applications such as the identification of topical words and information retrieval.
Year
DOI
Venue
2011
10.1088/1367-2630/13/12/123024
NEW JOURNAL OF PHYSICS
Keywords
Field
DocType
information retrieval,machine learning,shortest path,complex network,clustering coefficient,nearest neighbor,dynamic system
Zipf's law,Skewness,Shortest path problem,Quantum mechanics,Intermittency,Theoretical computer science,Betweenness centrality,Burstiness,Complex network,Clustering coefficient,Physics
Journal
Volume
Issue
ISSN
13
12
1367-2630
Citations 
PageRank 
References 
16
1.22
5
Authors
4
Name
Order
Citations
PageRank
Diego R. Amancio135229.53
Eduardo G. Altmann212412.01
Osvaldo N. Oliveira Jr.324717.25
Luciano da Fontoura Costa454263.09