A Multilingual Datasets Repository of the Hadith Content

Abstract

Knowledge extraction from unstructured data is a challenging research problem in research domain of Natural Language Processing (NLP). It requires complex NLP tasks like entity extraction and Information Extraction (IE), but one of the most challenging tasks is to extract all the required entities of data in the form of structured format so that data analysis can be applied. Our focus is to explain how the data is extracted in the form of datasets or conventional database so that further text and data analysis can be carried out. This paper presents a framework for Hadith data extraction from the Hadith authentic sources. Hadith is the collection of sayings of Holy Prophet Muhammad, who is the last holy prophet according to Islamic teachings. This paper discusses the preparation of the dataset repository and highlights issues in the relevant research domain. The research problem and their solutions of data extraction, pre-processing and data analysis are elaborated. The results have been evaluated using the standard performance evaluation measures. The dataset is available in multiple languages, multiple formats and is available free of cost for research purposes.

Authors and Affiliations

Ahsan Mahmood, Hikmat Ullah Khan, Fawaz K. Alarfaj, Muhammad Ramzan, Mahwish Ilyas

Keywords

Related Articles

Object’s Shape Recognition using Local Binary Patterns

This paper discusses the concept of object’s shape identification using local binary pattern technique (LBP). Since LBP is computationally simple it has been utilized successfully for recognition of various objects. LBP...

Privacy Preserving Data Publishing: A Classification Perspective

The concept of privacy is expressed as release of information in a controlled way. Privacy could also be defined as privacy decides what type of personal information should be released and which group or person can acces...

A hybrid Evolutionary Functional Link Artificial Neural Network for Data mining and Classification

This paper presents a specific structure of neural network as the functional link artificial neural network (FLANN). This technique has been employed for classification tasks of data mining. In fact, there are a few stud...

TGRP: A New Hybrid Grid-based Routing Approach for Manets

Most existing grid-based routing protocols use reactive mechanisms to build routing paths. In this paper, we propose a new hybrid approach for grid-based routing in MANETs which uses a combination of reactive and proacti...

Grid Approximation Based Inductive Charger Deployment Technique in Wireless Sensor Networks

Ensuring sufficient power in a sensor node is a challenging problem now-a-days to provide required level of security and data processing capability demanded by various applications scampered in a wireless sensor network....

Download PDF file
  • EP ID EP276758
  • DOI 10.14569/IJACSA.2018.090224
  • Views 103
  • Downloads 0

How To Cite

Ahsan Mahmood, Hikmat Ullah Khan, Fawaz K. Alarfaj, Muhammad Ramzan, Mahwish Ilyas (2018). A Multilingual Datasets Repository of the Hadith Content. International Journal of Advanced Computer Science & Applications, 9(2), 165-172. https://europub.co.uk/articles/-A-276758