Full-length isoform concatenation sequencing to resolve cancer transcriptome complexity

Saranga Wijeratne; Maria E. Hernandez Gonzalez; Kelli Roach; Katherine E. Miller; Kathleen M. Schieffer; James R. Fitch; Jeffrey Leonard; Peter White; Benjamin J. Kelly; Catherine E. Cottrell; Elaine R. Mardis; Richard K. Wilson; Anthony R. Miller

doi:10.1186/s12864-024-10021-x

BMC Genomics (Jan 2024)

Full-length isoform concatenation sequencing to resolve cancer transcriptome complexity

Saranga Wijeratne,
Maria E. Hernandez Gonzalez,
Kelli Roach,
Katherine E. Miller,
Kathleen M. Schieffer,
James R. Fitch,
Jeffrey Leonard,
Peter White,
Benjamin J. Kelly,
Catherine E. Cottrell,
Elaine R. Mardis,
Richard K. Wilson,
Anthony R. Miller

Affiliations

Saranga Wijeratne: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Maria E. Hernandez Gonzalez: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Kelli Roach: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Katherine E. Miller: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Kathleen M. Schieffer: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
James R. Fitch: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Jeffrey Leonard: Department of Neurosurgery, Nationwide Children’s Hospital
Peter White: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Benjamin J. Kelly: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Catherine E. Cottrell: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Elaine R. Mardis: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Richard K. Wilson: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital
Anthony R. Miller: The Steve and Cindy Rasmussen Institute for Genomic Medicine, Abigail Wexner Research Institute at Nationwide Children’s Hospital

DOI: https://doi.org/10.1186/s12864-024-10021-x
Journal volume & issue: Vol. 25, no. 1
pp. 1 – 19

Abstract

Read online

Abstract Background Cancers exhibit complex transcriptomes with aberrant splicing that induces isoform-level differential expression compared to non-diseased tissues. Transcriptomic profiling using short-read sequencing has utility in providing a cost-effective approach for evaluating isoform expression, although short-read assembly displays limitations in the accurate inference of full-length transcripts. Long-read RNA sequencing (Iso-Seq), using the Pacific Biosciences (PacBio) platform, can overcome such limitations by providing full-length isoform sequence resolution which requires no read assembly and represents native expressed transcripts. A constraint of the Iso-Seq protocol is due to fewer reads output per instrument run, which, as an example, can consequently affect the detection of lowly expressed transcripts. To address these deficiencies, we developed a concatenation workflow, PacBio Full-Length Isoform Concatemer Sequencing (PB_FLIC-Seq), designed to increase the number of unique, sequenced PacBio long-reads thereby improving overall detection of unique isoforms. In addition, we anticipate that the increase in read depth will help improve the detection of moderate to low-level expressed isoforms. Results In sequencing a commercial reference (Spike-In RNA Variants; SIRV) with known isoform complexity we demonstrated a 3.4-fold increase in read output per run and improved SIRV recall when using the PB_FLIC-Seq method compared to the same samples processed with the Iso-Seq protocol. We applied this protocol to a translational cancer case, also demonstrating the utility of the PB_FLIC-Seq method for identifying differential full-length isoform expression in a pediatric diffuse midline glioma compared to its adjacent non-malignant tissue. Our data analysis revealed increased expression of extracellular matrix (ECM) genes within the tumor sample, including an isoform of the Secreted Protein Acidic and Cysteine Rich (SPARC) gene that was expressed 11,676-fold higher than in the adjacent non-malignant tissue. Finally, by using the PB_FLIC-Seq method, we detected several cancer-specific novel isoforms. Conclusion This work describes a concatenation-based methodology for increasing the number of sequenced full-length isoform reads on the PacBio platform, yielding improved discovery of expressed isoforms. We applied this workflow to profile the transcriptome of a pediatric diffuse midline glioma and adjacent non-malignant tissue. Our findings of cancer-specific novel isoform expression further highlight the importance of long-read sequencing for characterization of complex tumor transcriptomes.

Published in BMC Genomics

ISSN: 1471-2164 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Technology: Chemical technology: Biotechnology; Science: Biology (General): Genetics
Website: http://bmcgenomics.biomedcentral.com

About the journal

Abstract

Keywords