On QSAR-based cardiotoxicity modeling with the expressiveness-enhanced graph learning model and dual-threshold scheme

Huijia Wang; Guangxian Zhu; Leighton T. Izu; Ye Chen-Izu; Naoaki Ono; MD Altaf-Ul-Amin; Shigehiko Kanaya; Ming Huang

doi:10.3389/fphys.2023.1156286

Frontiers in Physiology (May 2023)

On QSAR-based cardiotoxicity modeling with the expressiveness-enhanced graph learning model and dual-threshold scheme

Huijia Wang,
Guangxian Zhu,
Leighton T. Izu,
Ye Chen-Izu,
Naoaki Ono,
MD Altaf-Ul-Amin,
Shigehiko Kanaya,
Ming Huang

Affiliations

Huijia Wang: Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, Japan
Guangxian Zhu: Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, Japan
Leighton T. Izu: Department of Pharmacology, University of California, Davis, CA, United States
Ye Chen-Izu: Department of Biomedical Engineering, University of California, Davis, CA, United States
Naoaki Ono: Data Science Center, Nara Institute of Science and Technology, Ikoma, Japan
MD Altaf-Ul-Amin: Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, Japan
Shigehiko Kanaya: Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, Japan
Ming Huang: Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, Japan

DOI: https://doi.org/10.3389/fphys.2023.1156286
Journal volume & issue: Vol. 14

Abstract

Read online

Introduction: Given the direct association with malignant ventricular arrhythmias, cardiotoxicity is a major concern in drug design. In the past decades, computational models based on the quantitative structure–activity relationship have been proposed to screen out cardiotoxic compounds and have shown promising results. The combination of molecular fingerprint and the machine learning model shows stable performance for a wide spectrum of problems; however, not long after the advent of the graph neural network (GNN) deep learning model and its variant (e.g., graph transformer), it has become the principal way of quantitative structure–activity relationship-based modeling for its high flexibility in feature extraction and decision rule generation. Despite all these progresses, the expressiveness (the ability of a program to identify non-isomorphic graph structures) of the GNN model is bounded by the WL isomorphism test, and a suitable thresholding scheme that relates directly to the sensitivity and credibility of a model is still an open question.Methods: In this research, we further improved the expressiveness of the GNN model by introducing the substructure-aware bias by the graph subgraph transformer network model. Moreover, to propose the most appropriate thresholding scheme, a comprehensive comparison of the thresholding schemes was conducted.Results: Based on these improvements, the best model attains performance with 90.4% precision, 90.4% recall, and 90.5% F1-score with a dual-threshold scheme (active: <1μM; non-active: >30μM). The improved pipeline (graph subgraph transformer network model and thresholding scheme) also shows its advantages in terms of the activity cliff problem and model interpretability.

Published in Frontiers in Physiology

ISSN: 1664-042X (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Science: Physiology
Website: https://www.frontiersin.org/journals/physiology

About the journal

Abstract

Keywords