M.Phil / MS

Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/52

This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.

Browse

Search Results

Now showing 1 - 10 of 10
  • Item
    Graph Topological Indices and Machine Learning for Protein Networks and Drug Property Prediction
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2025) Ahmad Mehmood; CIIT/SP24-RMT-025/LHR; Dr. Sana Javed; LHR TP 10085
    Molecular property prediction for drugs is an integral part of computational chemistry in the early stages of drug discovery for efficient screening and optimization before experi- mental verification. In the present work, an exhaustive and interpretable machine learning (ML) strategy is proposed that combines SMILES-based molecular feature extraction with topological indices using molecular graph theoretical approaches for the prediction of key physicochemical properties of drugs. The dataset consisting of forty-seven pharmacolog- ically active molecules from DrugBank, with their structures characterized using sixteen topological descriptors such as the Wiener, Balaban, Harary, Randi´c, Zagreb, Schultz, and Shannon Entropy indexes, were identified using the RDKit grouping library coupled with mathematical formulations from molecular graph theory. Preprocessing steps were per- formed rigorously for missing values using imputation; removal with cutoffs using the inter-quartile range method; variance for stabilization using the Box-Cox transformation; and normalization with MinMax scaling. Four different regression models, namely Ordi- nary Least Squares (OLS), Ridge Regression, LASSO, and Elastic Net algorithms, were employed for the prediction of the MW, LogP, and HBD properties for the dataset with evaluations made in terms of R2, MAE, and Root MSE metrics; with feature interpreta- tion from SHapley Additive exPlanations (SHAP) analyses for feature interpretation. The results achieved for MW with high predictive precision with R2 ≈ 0.96 using linear mod- els effectively; moderate accuracy with R2 ≈ 0.4 for Lipophilicity; with poorer accuracy for HBD values likely due to their dependence on chemical principles underlying chem- ical groupings in chemistry. The interpretation using SHAP explained the contributions made by the Shannon entropy, Wiener, Zagreb, Schultz indices being the principal predic- ix tors in each case. The present work clearly illustrates the utility of integrating molecular chem-graph principles with interpretable ML algorithms for efficient, scalable, and insight- ful property-based predictions in early stage computational screens for drug identification strategies. Protein-protein interactions (PPIs) networks are very important to decode cellular signal- ing and mechanisms of disease. PPI also play a key to spot possible drug targets. A strong computational strategy is presented in this paper. It unites the graph theoretic approach with unsupervised machine learning. This study is used to detect biologically significant hub genes in a large scale lung cancer associated PPI network. The network retrieved from the STRING database (461 proteins, 18,704 weighted interactions). Using NetworkX, we constructed an undirected graph by extracting source and target node from the dataset. We computed 17 comprehensive topological centrality measures, confirming the networks scale free and small world properties (average clustering coefficient = 0.67, average short- est path length ≈ 1.96). The Isolation Forest algorithm identified 61 topological outliers. These outliers were related to nodes which had a very high influence. The K-means cluster- ing (that was optimized by the elbow curve) sorted the rest of the nodes into three groups. The hub genes were given ranks according to degree centrality and a composite centrality score. Among the genes, CALM3, CREB1, AKT1, MAPK1, EGFR and KRAS came out as the strongest candidates. The pathway enrichment analysis performed with KEGG and Reactome showed that the oncogenic pathways were remarkably over represented. These pathways were PI3KAkt, Ras, MAPK, and proteoglycans in cancer and EGFR signaling. Our methodology, which combines together the multiple centrality metrics, anomaly de- tection, and clustering, has indeed solved the drawbacks of traditional single metric ap- proaches. It has also provided greater sensitivity in hub detection. The hubs that were identified not only correspond to the well established lung cancer drivers but also point to potential novel biomarkers. The detected targets are thus a scalable and reproducible pipeline for systems level analysis of disease specific interactomes
  • Item
    A Study on Astragaloside IV in View of Connectivity Indices
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Muhammad Tahir Hussain; CIIT/SP20-RMT-003/LHR; Dr. Sana Javed; LHR TP 7622
    Topological indices are frequently used in the analysis of different chemical or molecular compounds including drugs. This thesis mainly deals with the computation of degree-based topological indices of the drug Astrgaloside IV and Benzenoid and the construction of a network of the indices. The main objective is to detect the subnetworks of the network with highly connected links called modules and the master regulators in that module. This approach would help the researchers to investigate the physio-chemical characteristics of a Astrgaloside IV based on the master regulatory index in the module where a master regulator index is an index that is at the top of a regulatory hierarchy in a module and is not influenced by any other index. This thesis is divided into five chapters. Chapter 1 includes the basic definitions , notions , and terminologies related to this thesis. Chapter 2 provides a review of literature about our work. The main results are gives in chapter 3. Chapter 4 contains a bried discussion about the work while chapter 5 consists of the references used in this thesis.
  • Item
    Molecular Topological Indices Based Analysis of Thermodynamics Properties of Terbium Dioxide
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Amir Hassan; CIIT/SP20-RMT-025/LHR; Dr. Sana Javed; LHR TP 7633
    The term used to illustrate a molecule/chemical compound in the form of graph is known as molecular/chemical graph. Molecules are usually represented as vertices while their bonding interaction is shown by edges in a molecular graph. In this thesis, we computed various connectivity indices based on degrees of vertices of chemical graph of Terbium Dioxide (Tb𝑂2) and Graphitic Carbon Nitride (g-𝐶3𝑁4) including general Randic, ABC, GA and Zagreb indices etc. Afterwards, we found the physical measures like entropy and heat of formation of Tb𝑂2 and g-𝐶3𝑁4. Then, we fitted curves between different indices and the thermodynamical properties namely heat of formation and entropy. Curve fitting was done in MATLAB through different methods based on linearity and non-linearity. The performance of the method was tested using root mean squared error (RMSE), the sum of squared errors (SSE) or R2. Further, we gave graphical representations of these indices. These mathematical frameworks might provide a way to study the ther - modynamics properties of the chemical structure of Terbium Dioxide (Tb𝑂2) at intense level which will assist to comprehend the relationship between system dimension and these measures. This thesis is divided into five chapters. Chapter 1 includes the basic definition , notions and terminologies related to thesis. Chapter 2 provides a review of literature about out work .The main results are given in chapter 3 .Chapter 4 contains a brief discussion about the work while the chapter 5 consists of the references used in this thesis
  • Item
    Describing Topology of Chemical Compounds Using Graphical Descriptors
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Mazhar Hussain; CIIT/SP20-RMT-004/LHR; Dr. Sana Javed; LHR TP 7623
    In terms of graph theory, a molecular graph or chemical graph is a representation of the structural formula of a chemical compound. A chemical graph is a labeled graph in which the vertices represent the atoms of a compound and the edges represent chemical bonds between two atoms. In this thesis, we computed various connectivity indices based on degrees of vertices of chemical graph of Iron(III) Oxide (𝐹𝑒2𝑂3) and Titanium Dioxide (𝑇𝑖𝑂2) including general Randic, ABC, GA and Zagreb indices etc. Afterwards, we found the physical measures like entropy and heat of formation of Iron(III) Oxide (𝐹𝑒2𝑂3) and Titanium Dioxide (𝑇𝑖𝑂2). Then, we fitted curves between different indices and the thermodynamical properties namely heat of formation and entropy. Curve fitting was done in MATLAB through different methods based on linearity and non-linearity. The performance of the method was tested using mean squared error (MSE), the sum of squared errors (SSE) or R2. Further, we gave graphical representations of these indices. These mathematical frameworks might provide a way to study the thermodynamical properties of the chemical structure Iron(III) Oxide (𝐹𝑒2𝑂3) and Titanium Dioxide (𝑇𝑖𝑂2) at intense level which will assist to comprehend the relationship between system dimension and these measures. This thesis is divided into five chapters. Chapter 1 contains basic definitions, notions and terminologies related to this thesis. Chapter 2 provides a review of literature about our work. The main results are given in chapter 3. Chapter 4 contains a brief discussion about the work while chapter 5 consist of the references used in this thesis.
  • Item
    Molecular Interpretation Of Indium Phosphate Using
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Waqar Asghar; FA19-RMT-061; Dr. Sana Javed; LHR TP 7403
    This study presents a molecular interpretation of indium phosphate using chemical graph theory and structural modeling approaches. Indium phosphate is analyzed at the molecular level by representing its atomic structure as a graph, where vertices correspond to atoms and edges correspond to chemical bonds. The research focuses on understanding the structural arrangement, bonding characteristics, and topological properties of the compound. Various molecular descriptors and topological indices are computed to capture the physicochemical behavior of the material. The study highlights how graph-theoretical techniques can be used to interpret the stability, reactivity, and electronic properties of indium phosphate, providing a mathematical framework for analyzing complex inorganic materials.
  • Item
    Mathematical Analysis of Delayed Tumor Models
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2018) Maria Sajjad; CIIT/SP17-RMT-030/LHR; Dr. Sana Javed; LHR TP 5463
    This thesis focuses on the current problem of cancer, interlinked with innate immune system. We have taken into account the innate immune system with delay and the delay differential equations are solved using analytical and nu- merical techniques. From the mathematical analysis we have explored the chal- lenging problem of cancer invasion.
  • Item
    Emerging Patterns in Machine Learning for Disease Detection
    (Library Information Services COMSATS University Lahore Campus, 2024-03-18) Shagufta Mufeez; SP23-RMT-038; Dr. Sana Javed; LHR TP 9582
    The growing number of diseases and disorders worldwide further supports the need for prompt and precise diagnosis and classification. Understanding how disease diagnosis can be enhanced through the use of reliable computational techniques including pre process ing approaches, feature selection methods, and hyper parameter tuning with Grid Search Cross-Validation (Grid Search CV), primarily boosts and optimizes feature selection from the machine learning field is the goal of this thesis.” These procedures are essential for ad dressing problems with the developed models’ performance, interpretability, and depend ability. The dataset in the context of disease classification is a comprehensive collection of mea sures. Missing values and encoding have been addressed at the pre processing stage of the data handling process to guarantee a stable and clean baseline for the subsequent steps. Feature selection techniques are employed in order to eliminate insignificant instances and prevent the issue of their dominance. Decision Tree, Random Forest, Support Vector Classifier (SVC), XG Boost, Light GBM and CAT Boost are the six machine learning models that were applied to the dataset. The Grid Search Cross Validation process, which enables the proper configuration of hyper pa rameters, is used to tune the parameters of each model. The evaluation metrics such as confusion matrix, accuracy, precision, recall and F1 score are used to assess models’ per formance. These findings demonstrate how feature selection and hyperparameter modification may be used to enhance the model. This is an interesting finding that raises the possibility of using machine learning to increase early diagnosis rates and decrease diagnostic errors. This study advances the field of machine learning in healthcare and offers a method for effectively and efficiently detecting illnesses.
  • Item
    Machine Learning Methodologies for Forecasting Drug Properties and Identifying Disease Indicators
    (Library Information Services COMSATS University Lahore Campus, 2024-03) Muhammad Owais; FA22-RMT-015; Dr. Sana Javed; LHR TP 9385
    This thesis investigates the use of machine learning approaches to predict and analyze drug properties in terms of topological indices, which are important for understanding their chemical and biological characteristics. Topological indices, generated from graphical rep resentation of chemical formation of a drug, give a measurable assessment of the molecule’s structure and are widely utilized in drug design and discovery. This study uses advanced machine learning methods to improve the accuracy and efficiency of predicting the drug property namely molecular weight, allowing for faster screening and optimization of med ication candidates. The number of patients registered to the hospitals diagnosed with liver disorder is very high. ML approaches might be utilized to overcome the burden on the doctors by developing accurate classifiers for disease prediction. It is also advantageous to detect the key factors involved in the development of the disease so that precautionary measures might be taken for prevention. This thesis is also focused to develop a machine learning classifier to predict liver disease.
  • Item
    Chromatic Equivalence and Uniqueness of Subdivided K4 Graphs
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2016) Areeb; FA14-MSMATH-009; LHR TP 6935; Dr. Sana Javed
    In this thesis, various chromatically unique and correspondent pairs of several families of graphs with same girths have been attained. Subdivided complete graphs are related to these families, indeed. Our major goal in this paper is to carry on the study of chromaticity of the subdivided complete graphs. We are dealing particularly with four vertices which is called 𝑘4-homeomorph. Distinctiveness of two particular classes of 𝑘4-homeomorph graphs with girth 11 has been explained in detail. The cycles of different lengths from graph has been removed during the discussion. In order to make them chromatically unique, vital and satisfactory conditions have been established. All well known definitions, results and terms are given in Chapter 1. Chapter 2 is covering a brief literature review about chromatic uniqueness and correspondence. While the chapter 3 is focused on the chromatic uniqueness along with chromatic equivalence of two 𝐾4-homeomorph graphs with same girth. At the end of the thesis, few open problems have been recommended as well.
  • Item
    Performance Analysis of Bayesian Learning Against Other Conventional Machine Learning Techniques
    (Library Information Services COMSATS University Islamabad Lahore Campus, 2025) Hassan Aftab (FA23-RMT-054); Dr. Sana Javed; LHR TP 9797
    This study comprises of performance analysis of variants of Naive Bayes and other conventional machine learning algorithms. The mathematics and core intuition behind the machine learning algorithms has been discussed in this text. Models have been trained and tested on different datasets having different genre of features such as continuous, categorical and discrete. Performance analysis has been done to see which model performs better on which dataset. Feature engineering techniques have also been employed and analysis has been done to infer how they impact model’s performance. Breast cancer dataset, heart attack dataset and air quality dataset have been used for the purpose of analysis of results of models on data having continuous independent features. Customized versions of the Naive Bayes algorithms have also been developed by taking the likelihood of features from the probability density functions of Lognorm, Student’s t and Skewnorm distributions. For analysis on data having categorical input features, datasets of Tic-Tac-Toe game, Car evaluation and Connect-4 Game have been used. As far as the discrete features are concerned, Fashion MNIST dataset having pixel values of images as features has been used. An email spam classifie
2025 @ COMSATS University Islamabad, Lahore Campus. All rights reserved.