A Machine Learning Approach for Mutation Identification in Gene Sequences to Detect Lungs Cancer
No Thumbnail Available
Date
2022
Journal Title
Journal ISSN
Volume Title
Publisher
Library Information Services, COMSATS University Islamabad, Lahore Campus
Abstract
Cancer is one of the leading causes of death worldwide. The mutation is one of the basic causes of cancer, which is caused by any change in the gene sequences of the human body. The abnormal and uncontrolled growth in the cells of the lungs is a cause of lung cancer. It is the second most common type of cancer worldwide, considered 11.2% of all cancers. According to the cancer statistics of the United States of America (USA), about 236,740 people were diagnosed with lung cancer in the year 2021, and about 130,180 people died from lung cancer in the same year. This study aims to develop a dataset, machine learning, and deep learning-based techniques to address the above-mentioned issues. In this regard, the asia.ensembl.org website is used to extract normal gene sequences from the lung cancer dataset, while the mutated gene sequences for lung cancer are obtained from the intogen.org website. Thus, the dataset of 1588 samples consisting of 2576 mutations is developed by following the standard dataset creation method. Then, different feature extraction methods are applied to extract the discriminative features from the dataset; these features include raw moment, Hahn moment, central moment, position relative incident matrix (PRIM), reverse position relative incident matrix (RPRIM) calculation, feature vector determination, accumulative absolute position incidence vector (AAPIV), reverse accumulative absolute position incidence vector (RAAPIV). These features are then fed to the traditional machine learning (ML) algorithms such as Random Forest (RF), Logistic Regression (LR), Multi-Layer Perceptron (MLP), and an ensemble of these traditional ML algorithms. These extracted features are also fed to deep learning (DL) methods such as LSTM, GRU, Bi-LSTM, and an ensemble of these DL algorithms. An averaging technique is applied to ensemble the classification algorithms. Three validation methods are used for evaluation: self consistency testing, 10-fold cross-validation, and independent set testing. The final classification result of lung cancer is generated as accuracy, specificity, Mathew's Correlation Coefficient (MCC), sensitivity, and Cohen's Kappa. The efficacy of this study can be measured with the accuracy of 94.5% obtained from the ensemble of x machine learning models using 10-FCV and with the accuracy of 97.3% obtained from the ensemble of deep learning models using the independent set test.
Description
Keywords
Dr. Allah Bux Sargano, Sp19, Department of Computer Science, TECHNOLOGY::Information technology::Computer science, Detect Lungs Cancer