Title
Conjugation-based compression for Hebrew texts
Abstract
Traditional compression techniques do not look deeply into the morphology of languages. This can be less critical in languages like English where most of the sequences are illegal according to the grammatical rules of the language, for example, zx, bv or qe; hence the morphology can add a little information that can be beneficial for the compression algorithm. However, this negligence can be a significant flaw in languages like Hebrew where the grammatical rules allow much more freedom in the sequences of letters and, except tet after gimel, any pair is legal; hence compressing without taking the morphological rules into account can yield a poorer compression ratio. This article suggests a tool that optimizes the Burrows-Wheeler algorithm which is an unaware morphological rules compression method. It first preprocesses a Hebrew text file according to the Hebrew conjugation rules, and, after that, it provides the Burrows-Wheeler algorithm with this preprocessed file so that can be compressed better. Experimental results show a significant improvement.
Year
DOI
Venue
2007
10.1145/1227850.1227854
ACM Trans. Asian Lang. Inf. Process.
Keywords
Field
DocType
morphological rule,compression algorithm,hebrew conjugation rule,traditional compression technique,hebrew text file,conjugation-based compression,preprocessed file,root conjugations,languages additional key words and phrases: text compression,semitic languages acm reference format:,hebrew text analysis,burrows-wheeler algorithm,compression method,grammatical rule,categories and subject descriptors: e.4 coding and information theory: data completion and compression general terms: algorithms,poorer compression ratio,information theory,compression ratio,semitic languages,text analysis
Compression (physics),Text compression,Computer science,Semitic languages,Hebrew,Speech recognition,Compression ratio,Artificial intelligence,Natural language processing,Data compression
Journal
Volume
Issue
Citations 
6
1
2
PageRank 
References 
Authors
0.40
16
2
Name
Order
Citations
PageRank
Yair Wiseman115814.60
Irit Gefner220.40