Biologically meaningful genome interpretation models to address data underdetermination for the leaf and seed ionome prediction in Arabidopsis thaliana

Daniele Raimondi; Antoine Passemiers; Nora Verplaetse; Massimiliano Corso; Ángel Ferrero-Serrano; Nelson Nazzicari; Filippo Biscarini; Piero Fariselli; Yves Moreau

doi:10.1038/s41598-024-63855-6

Scientific Reports (Jun 2024)

Biologically meaningful genome interpretation models to address data underdetermination for the leaf and seed ionome prediction in Arabidopsis thaliana

Daniele Raimondi,
Antoine Passemiers,
Nora Verplaetse,
Massimiliano Corso,
Ángel Ferrero-Serrano,
Nelson Nazzicari,
Filippo Biscarini,
Piero Fariselli,
Yves Moreau

Affiliations

Daniele Raimondi: ESAT-STADIUS, KU Leuven
Antoine Passemiers: ESAT-STADIUS, KU Leuven
Nora Verplaetse: ESAT-STADIUS, KU Leuven
Massimiliano Corso: Université Paris-Saclay, INRAE, AgroParisTech, Institute Jean-Pierre Bourgin for Plant Sciences (IJPB)
Ángel Ferrero-Serrano: Department of Biology, Pennsylvania State University
Nelson Nazzicari: CREA-ZA
Filippo Biscarini: CNR-IBBA
Piero Fariselli: Department of Medical Sciences, University of Torino
Yves Moreau: ESAT-STADIUS, KU Leuven

DOI: https://doi.org/10.1038/s41598-024-63855-6
Journal volume & issue: Vol. 14, no. 1
pp. 1 – 11

Abstract

Read online

Abstract Genome interpretation (GI) encompasses the computational attempts to model the relationship between genotype and phenotype with the goal of understanding how the first leads to the second. While traditional approaches have focused on sub-problems such as predicting the effect of single nucleotide variants or finding genetic associations, recent advances in neural networks (NNs) have made it possible to develop end-to-end GI models that take genomic data as input and predict phenotypes as output. However, technical and modeling issues still need to be fixed for these models to be effective, including the widespread underdetermination of genomic datasets, making them unsuitable for training large, overfitting-prone, NNs. Here we propose novel GI models to address this issue, exploring the use of two types of transfer learning approaches and proposing a novel Biologically Meaningful Sparse NN layer specifically designed for end-to-end GI. Our models predict the leaf and seed ionome in A.thaliana, obtaining comparable results to our previous over-parameterized model while reducing the number of parameters by 8.8 folds. We also investigate how the effect of population stratification influences the evaluation of the performances, highlighting how it leads to (1) an instance of the Simpson’s Paradox, and (2) model generalization limitations.

Published in Scientific Reports

ISSN: 2045-2322 (Online)
Publisher: Nature Portfolio
Country of publisher: United Kingdom
LCC subjects: Medicine; Science
Website: https://www.nature.com/srep/

About the journal

Abstract

Keywords