Genome Biology (Oct 2022)

iDNA-ABF: multi-scale deep biological language learning model for the interpretable prediction of DNA methylations

  • Junru Jin,
  • Yingying Yu,
  • Ruheng Wang,
  • Xin Zeng,
  • Chao Pang,
  • Yi Jiang,
  • Zhongshen Li,
  • Yutong Dai,
  • Ran Su,
  • Quan Zou,
  • Kenta Nakai,
  • Leyi Wei

DOI
https://doi.org/10.1186/s13059-022-02780-1
Journal volume & issue
Vol. 23, no. 1
pp. 1 – 23

Abstract

Read online

Abstract In this study, we propose iDNA-ABF, a multi-scale deep biological language learning model that enables the interpretable prediction of DNA methylations based on genomic sequences only. Benchmarking comparisons show that our iDNA-ABF outperforms state-of-the-art methods for different methylation predictions. Importantly, we show the power of deep language learning in capturing both sequential and functional semantics information from background genomes. Moreover, by integrating the interpretable analysis mechanism, we well explain what the model learns, helping us build the mapping from the discovery of important sequential determinants to the in-depth analysis of their biological functions.

Keywords