Title
Natural Language Processing For Web Browsing Analytics: Challenges, Lessons Learned, And Opportunities
Abstract
In an Internet arena where the search engines and other digital marketing firms' revenues peak, other actors still have open opportunities to monetize their users' data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks' links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF-IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrization.
Year
DOI
Venue
2021
10.1016/j.comnet.2021.108357
COMPUTER NETWORKS
Keywords
DocType
Volume
Web browsing, Users analytics, Natural language processing, Deep learning, Traffic monetization, Internet monitoring
Journal
198
ISSN
Citations 
PageRank 
1389-1286
0
0.34
References 
Authors
0
5
Name
Order
Citations
PageRank
Daniel Perdices101.01
Javier Ramos2488.30
José Luis García-Dorado39513.01
Ivan Gonzalez411016.67
Jorge E. López de Vergara518726.98