Deep Learning Classification of Biomedical Text using Convolutional Neural Network
Journal Title: International Journal of Advanced Computer Science & Applications - Year 2019, Vol 10, Issue 8
Abstract
In this digital era, the document entries have been increasing days by days, causing a situation where the volume of the document entries in overwhelming. This situation has caused people to encounter with problems such as congestion of data, difficulty in searching the intended information or even difficulty in managing the databases, for example, MEDLINE database which stores the documents related to the biomedical field. This research will specify the solution focusing in text classification of the biomedical abstracts. Text classification is the process of organizing documents into predefined classes. A standard text classification framework consists of feature extraction, feature selection and the classification stages. The dataset used in this research is the Ohsumed dataset which is the subset of the MEDLINE database. In this research, there is a total number of 11,566 abstracts selected from the Ohsumed dataset. First of all, feature extraction is performed on the biomedical abstracts and a list of unique features is produced. All the features in this list will be added to the multiword tokenizer lexicon for tokenizing phrases or compound word. After that, the classification of the biomedical texts is conducted using the deep learning network, Convolutional Neural Network which is an approach widely used in many domains such as pattern recognition, classification and so on. The goal of classification is to accurately organize the data into the correct predefined classes. The Convolutional Neural Network has achieved a result of 54.79% average accuracy, 61.00% average precision, 60.00% average recall and 60.50% average F1-score. In short, it is hoped that this research could be beneficial to the text classification area.
Authors and Affiliations
Rozilawati Dollah, Chew Yi Sheng, Norhawaniah Zakaria, Mohd Shahizan Othman, Abd Wahid Rasib
A P System for K-Medoids-Based Clustering
The membrane computing model, also known as the P system, is a parallel and distributed computing system. K-medoids algorithm is one of the most famous algorithms in partition-based clustering algorithms, and has been wi...
An Improved Image Steganography Method Based on LSB Technique with Random Pixel Selection
With the rapid advance in digital network, information technology, digital libraries, and particularly World Wide Web services, many kinds of information could be retrieved any time. Thus, the security issue has become o...
Modeling and Implementing Ontology for Managing Learners’ Profiles
This paper presents an issue that is important to consider when developing a learning environment whose field is constantly evolving mainly in terms of the use of training platforms. Research in this field has enabled th...
A Novel Approach for Dimensionality Reduction and Classification of Hyperspectral Images based on Normalized Synergy
During the last decade, hyperspectral images have attracted increasing interest from researchers worldwide. They provide more detailed information about an observed area and allow an accurate target detection and precise...
A QUADRATIC CONVERGENCE METHOD FOR THE MANAGEMENT EQUILIBRIUM MODEL
In this paper, we study a class of methods for solving the management equilibrium model. We first give an estimate of the error bound for the model, and then, based on the estimate of the error bound, propose a method fo...