Graph Topological Indices and Machine Learning for Protein Networks and Drug Property Prediction

No Thumbnail Available

Date

2025

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus

Abstract

Molecular property prediction for drugs is an integral part of computational chemistry in the early stages of drug discovery for efficient screening and optimization before experi- mental verification. In the present work, an exhaustive and interpretable machine learning (ML) strategy is proposed that combines SMILES-based molecular feature extraction with topological indices using molecular graph theoretical approaches for the prediction of key physicochemical properties of drugs. The dataset consisting of forty-seven pharmacolog- ically active molecules from DrugBank, with their structures characterized using sixteen topological descriptors such as the Wiener, Balaban, Harary, Randi´c, Zagreb, Schultz, and Shannon Entropy indexes, were identified using the RDKit grouping library coupled with mathematical formulations from molecular graph theory. Preprocessing steps were per- formed rigorously for missing values using imputation; removal with cutoffs using the inter-quartile range method; variance for stabilization using the Box-Cox transformation; and normalization with MinMax scaling. Four different regression models, namely Ordi- nary Least Squares (OLS), Ridge Regression, LASSO, and Elastic Net algorithms, were employed for the prediction of the MW, LogP, and HBD properties for the dataset with evaluations made in terms of R2, MAE, and Root MSE metrics; with feature interpreta- tion from SHapley Additive exPlanations (SHAP) analyses for feature interpretation. The results achieved for MW with high predictive precision with R2 ≈ 0.96 using linear mod- els effectively; moderate accuracy with R2 ≈ 0.4 for Lipophilicity; with poorer accuracy for HBD values likely due to their dependence on chemical principles underlying chem- ical groupings in chemistry. The interpretation using SHAP explained the contributions made by the Shannon entropy, Wiener, Zagreb, Schultz indices being the principal predic- ix tors in each case. The present work clearly illustrates the utility of integrating molecular chem-graph principles with interpretable ML algorithms for efficient, scalable, and insight- ful property-based predictions in early stage computational screens for drug identification strategies. Protein-protein interactions (PPIs) networks are very important to decode cellular signal- ing and mechanisms of disease. PPI also play a key to spot possible drug targets. A strong computational strategy is presented in this paper. It unites the graph theoretic approach with unsupervised machine learning. This study is used to detect biologically significant hub genes in a large scale lung cancer associated PPI network. The network retrieved from the STRING database (461 proteins, 18,704 weighted interactions). Using NetworkX, we constructed an undirected graph by extracting source and target node from the dataset. We computed 17 comprehensive topological centrality measures, confirming the networks scale free and small world properties (average clustering coefficient = 0.67, average short- est path length ≈ 1.96). The Isolation Forest algorithm identified 61 topological outliers. These outliers were related to nodes which had a very high influence. The K-means cluster- ing (that was optimized by the elbow curve) sorted the rest of the nodes into three groups. The hub genes were given ranks according to degree centrality and a composite centrality score. Among the genes, CALM3, CREB1, AKT1, MAPK1, EGFR and KRAS came out as the strongest candidates. The pathway enrichment analysis performed with KEGG and Reactome showed that the oncogenic pathways were remarkably over represented. These pathways were PI3KAkt, Ras, MAPK, and proteoglycans in cancer and EGFR signaling. Our methodology, which combines together the multiple centrality metrics, anomaly de- tection, and clustering, has indeed solved the drawbacks of traditional single metric ap- proaches. It has also provided greater sensitivity in hub detection. The hubs that were identified not only correspond to the well established lung cancer drivers but also point to potential novel biomarkers. The detected targets are thus a scalable and reproducible pipeline for systems level analysis of disease specific interactomes

Description

Keywords

Department of Mathematics, SP24, Mathematics, Graph Topological Indices, Property Prediction, Machine Learning, Dr. Sana Javed

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By