OLAWSDS:An Online Arabic Web Spam Detection System
Journal Title: International Journal of Advanced Computer Science & Applications - Year 2014, Vol 5, Issue 2
Abstract
For marketing purposes, Some Websites designers and administrators use illegal Search Engine Optimization (SEO) techniques to optimize the ranking of their Web pages and mislead the search engines. Some Arabic Web pages use both content and link features, to increase artificially the rank of their Web pages in the Search Engine Results Pages (SERPs). This study represents an enhancement to previous work in this field. It includes the design and implementation of an online Arabic Web spam detection system, based on algorithms and mathematical foundations, which can detect the Arabic content and link web spam depending on the tree of the spam detection conditions, beside depending on the user’s feedback through a custom Web browser. The users can participate in making the decision about any Web page, through their feedbacks, so they judge if the Arabic Web pages in the browser are relevant for their particular queries or not. The proposed system uses the extracted content and link features from Arabic Web pages to determine whether to label each Web page as a spam or as a non-spam. This system also attempts to learn from the user’s feedback to enhance automatically its performance. Statistical analysis is adopted in this study to evaluate the proposed system. Statistical Package for the Social Sciences (SPSS) software is used to evaluate this new system which considers the users feedbacks as dependent variables, while Arabic content and links features on the other hand are considered independent variables. The statistical analysis with the SPSS is used to apply a variety of tests, such as the test of the analysis of variance (ANOVA). ANOVA is used to show the relationships between the dependent and independent variables in the dataset, which leads to solving problems and building intelligent decisions and results.
Authors and Affiliations
Mohammed Al-Kabi, Heider Wahsheh, Izzat Alsmadi
A New Selection Operator - CSM in Genetic Algorithms for Solving the TSP
Genetic Algorithms (GAs) is a type of local search that mimics biological evolution by taking a population of string, which encodes possible solutions and combines them based on fitness values to produce individuals that...
An Enhanced Concept based Approach for user Centered Health Information Retrieval to Address Readability Issues
Searching for relevant medical guidance has turn out to be a general and notable task executed by internet users. This diversity of quantifiable information explorers indicates the enormous range of information needs and...
Interest Reduction and PIT Minimization in Content Centric Networks
Content Centric Networking aspires to a more efficient use of the Internet through in-path caching, multi-homing, and provisions for state maintenance and intelligent forwarding at the CCN routers. However, these benefit...
Comparison of Digital Signature Algorithm and Authentication Schemes for H.264 Compressed Video
In this paper we present the advantages of the elliptic curve cryptography for the implementations of the electronic signature algorithms “elliptic curve digital signature algorithm, ECDSA”, compared with “the digital si...
Analysis and Maximizing Energy Harvesting from RF Signals using T-Shaped Microstrip Patch Antenna
The advancement of the modern world requires catering the power crisis. New methodologies for energy harvesting were considered, but their succession in a different environment is still to explore. This paper deals with...