Graph Topological Indices and Machine Learning for Protein Networks and Drug Property Prediction

dc.contributor.authorAhmad Mehmood
dc.contributor.authorCIIT/SP24-RMT-025/LHR
dc.contributor.authorDr. Sana Javed
dc.contributor.authorLHR TP 10085
dc.date.accessioned2026-06-01T13:39:07Z
dc.date.issued2025
dc.description.abstractMolecular property prediction for drugs is an integral part of computational chemistry in the early stages of drug discovery for efficient screening and optimization before experi- mental verification. In the present work, an exhaustive and interpretable machine learning (ML) strategy is proposed that combines SMILES-based molecular feature extraction with topological indices using molecular graph theoretical approaches for the prediction of key physicochemical properties of drugs. The dataset consisting of forty-seven pharmacolog- ically active molecules from DrugBank, with their structures characterized using sixteen topological descriptors such as the Wiener, Balaban, Harary, Randi´c, Zagreb, Schultz, and Shannon Entropy indexes, were identified using the RDKit grouping library coupled with mathematical formulations from molecular graph theory. Preprocessing steps were per- formed rigorously for missing values using imputation; removal with cutoffs using the inter-quartile range method; variance for stabilization using the Box-Cox transformation; and normalization with MinMax scaling. Four different regression models, namely Ordi- nary Least Squares (OLS), Ridge Regression, LASSO, and Elastic Net algorithms, were employed for the prediction of the MW, LogP, and HBD properties for the dataset with evaluations made in terms of R2, MAE, and Root MSE metrics; with feature interpreta- tion from SHapley Additive exPlanations (SHAP) analyses for feature interpretation. The results achieved for MW with high predictive precision with R2 ≈ 0.96 using linear mod- els effectively; moderate accuracy with R2 ≈ 0.4 for Lipophilicity; with poorer accuracy for HBD values likely due to their dependence on chemical principles underlying chem- ical groupings in chemistry. The interpretation using SHAP explained the contributions made by the Shannon entropy, Wiener, Zagreb, Schultz indices being the principal predic- ix tors in each case. The present work clearly illustrates the utility of integrating molecular chem-graph principles with interpretable ML algorithms for efficient, scalable, and insight- ful property-based predictions in early stage computational screens for drug identification strategies. Protein-protein interactions (PPIs) networks are very important to decode cellular signal- ing and mechanisms of disease. PPI also play a key to spot possible drug targets. A strong computational strategy is presented in this paper. It unites the graph theoretic approach with unsupervised machine learning. This study is used to detect biologically significant hub genes in a large scale lung cancer associated PPI network. The network retrieved from the STRING database (461 proteins, 18,704 weighted interactions). Using NetworkX, we constructed an undirected graph by extracting source and target node from the dataset. We computed 17 comprehensive topological centrality measures, confirming the networks scale free and small world properties (average clustering coefficient = 0.67, average short- est path length ≈ 1.96). The Isolation Forest algorithm identified 61 topological outliers. These outliers were related to nodes which had a very high influence. The K-means cluster- ing (that was optimized by the elbow curve) sorted the rest of the nodes into three groups. The hub genes were given ranks according to degree centrality and a composite centrality score. Among the genes, CALM3, CREB1, AKT1, MAPK1, EGFR and KRAS came out as the strongest candidates. The pathway enrichment analysis performed with KEGG and Reactome showed that the oncogenic pathways were remarkably over represented. These pathways were PI3KAkt, Ras, MAPK, and proteoglycans in cancer and EGFR signaling. Our methodology, which combines together the multiple centrality metrics, anomaly de- tection, and clustering, has indeed solved the drawbacks of traditional single metric ap- proaches. It has also provided greater sensitivity in hub detection. The hubs that were identified not only correspond to the well established lung cancer drivers but also point to potential novel biomarkers. The detected targets are thus a scalable and reproducible pipeline for systems level analysis of disease specific interactomes
dc.identifier.urihttps://repository.cuilahore.edu.pk/123456789/3980
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 10085
dc.subjectDepartment of Mathematics
dc.subjectSP24
dc.subjectMathematics
dc.subjectGraph Topological Indices
dc.subjectProperty Prediction
dc.subjectMachine Learning
dc.subjectDr. Sana Javed
dc.titleGraph Topological Indices and Machine Learning for Protein Networks and Drug Property Prediction
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Ahmad_Thesis.pdf
Size:
14.8 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections