Title
A Lexicon Based Approach For Classifying Arabic Multi-Labeled Text
Abstract
Purpose - Multi-label Text Classification (MTC) is one of the most recent research trends in data mining and information retrieval domains because of many reasons such as the rapid growth of online data and the increasing tendency of internet users to be more comfortable with assigning multiple labels/tags to describe documents, emails, posts, etc. The dimensionality of labels makes MTC more difficult and challenging compared with traditional single-labeled text classification (TC). Because it is a natural extension of TC, several ways are proposed to benefit from the rich literature of TC through what is called problem transformation (PT) methods. Basically, PT methods transform the multi-label data into a single-label one that is suitable for traditional single-label classification algorithms. Another approach is to design novel classification algorithms customized for MTC. Over the past decade, several works have appeared on both approaches focusing mainly on the English language. This work aims to present an elaborate study of MTC of Arabic articles.Design/methodology/approach - This paper presents a novel lexicon-based method for MTC, where the keywords that are most associated with each label are extracted from the training data along with a threshold that can later be used to determine whether each test document belongs to a certain label.Findings - The experiments show that the presented approach outperforms the currently available approaches. Specifically, the results of our experiments show that the best accuracy obtained from existing approaches is only 18 per cent, whereas the accuracy of the presented lexicon-based approach can reach an accuracy level of 31 per cent.Originality/value - Although there exist some tools that can be customized to address the MTC problem for Arabic text, their accuracies are very low when applied to Arabic articles. This paper presents a novel method for MTC. The experiments show that the presented approach outperforms the currently available approaches.
Year
DOI
Venue
2016
10.1108/IJWIS-01-2016-0002
INTERNATIONAL JOURNAL OF WEB INFORMATION SYSTEMS
Keywords
Field
DocType
Label-set dimensionality, Lexicon-based multi-label classification, ML-Accuracy, Multi-label data, Single-label data
Training set,Data mining,English language,Arabic,Computer science,Originality,Curse of dimensionality,Lexicon,Statistical classification,The Internet
Journal
Volume
Issue
ISSN
12
4
1744-0084
Citations 
PageRank 
References 
1
0.39
0
Authors
4
Name
Order
Citations
PageRank
Ismail Hmeidi19511.46
Mahmoud Al-Ayyoub273063.41
Nizar A. Mahyoub310.39
Mohammed A. Shehab41046.94