Graph Topological Indices and Machine Learning for Protein Networks and Drug Property Prediction
No Thumbnail Available
Date
2025
Journal Title
Journal ISSN
Volume Title
Publisher
Library Information Services, COMSATS University Islamabad, Lahore Campus
Abstract
Molecular property prediction for drugs is an integral part of computational chemistry in
the early stages of drug discovery for efficient screening and optimization before experi-
mental verification. In the present work, an exhaustive and interpretable machine learning
(ML) strategy is proposed that combines SMILES-based molecular feature extraction with
topological indices using molecular graph theoretical approaches for the prediction of key
physicochemical properties of drugs. The dataset consisting of forty-seven pharmacolog-
ically active molecules from DrugBank, with their structures characterized using sixteen
topological descriptors such as the Wiener, Balaban, Harary, Randi´c, Zagreb, Schultz, and
Shannon Entropy indexes, were identified using the RDKit grouping library coupled with
mathematical formulations from molecular graph theory. Preprocessing steps were per-
formed rigorously for missing values using imputation; removal with cutoffs using the
inter-quartile range method; variance for stabilization using the Box-Cox transformation;
and normalization with MinMax scaling. Four different regression models, namely Ordi-
nary Least Squares (OLS), Ridge Regression, LASSO, and Elastic Net algorithms, were
employed for the prediction of the MW, LogP, and HBD properties for the dataset with
evaluations made in terms of R2, MAE, and Root MSE metrics; with feature interpreta-
tion from SHapley Additive exPlanations (SHAP) analyses for feature interpretation. The
results achieved for MW with high predictive precision with R2 ≈ 0.96 using linear mod-
els effectively; moderate accuracy with R2 ≈ 0.4 for Lipophilicity; with poorer accuracy
for HBD values likely due to their dependence on chemical principles underlying chem-
ical groupings in chemistry. The interpretation using SHAP explained the contributions
made by the Shannon entropy, Wiener, Zagreb, Schultz indices being the principal predic-
ix
tors in each case. The present work clearly illustrates the utility of integrating molecular
chem-graph principles with interpretable ML algorithms for efficient, scalable, and insight-
ful property-based predictions in early stage computational screens for drug identification
strategies.
Protein-protein interactions (PPIs) networks are very important to decode cellular signal-
ing and mechanisms of disease. PPI also play a key to spot possible drug targets. A strong
computational strategy is presented in this paper. It unites the graph theoretic approach
with unsupervised machine learning. This study is used to detect biologically significant
hub genes in a large scale lung cancer associated PPI network. The network retrieved from
the STRING database (461 proteins, 18,704 weighted interactions). Using NetworkX, we
constructed an undirected graph by extracting source and target node from the dataset.
We computed 17 comprehensive topological centrality measures, confirming the networks
scale free and small world properties (average clustering coefficient = 0.67, average short-
est path length ≈ 1.96). The Isolation Forest algorithm identified 61 topological outliers.
These outliers were related to nodes which had a very high influence. The K-means cluster-
ing (that was optimized by the elbow curve) sorted the rest of the nodes into three groups.
The hub genes were given ranks according to degree centrality and a composite centrality
score. Among the genes, CALM3, CREB1, AKT1, MAPK1, EGFR and KRAS came out
as the strongest candidates. The pathway enrichment analysis performed with KEGG and
Reactome showed that the oncogenic pathways were remarkably over represented. These
pathways were PI3KAkt, Ras, MAPK, and proteoglycans in cancer and EGFR signaling.
Our methodology, which combines together the multiple centrality metrics, anomaly de-
tection, and clustering, has indeed solved the drawbacks of traditional single metric ap-
proaches. It has also provided greater sensitivity in hub detection. The hubs that were
identified not only correspond to the well established lung cancer drivers but also point
to potential novel biomarkers. The detected targets are thus a scalable and reproducible
pipeline for systems level analysis of disease specific interactomes
Description
Keywords
Department of Mathematics, SP24, Mathematics, Graph Topological Indices, Property Prediction, Machine Learning, Dr. Sana Javed