Author
Listed:
- Paul R Munn
- Jay Chia
- Charles G Danko
Abstract
Functional element annotations are critical tools used to provide insight into the molecular processes governing cell development, differentiation, and disease. Run-on and sequencing assays measure the production of nascent RNAs and can provide an effective data source for discovering functional elements. However, the accurate inference of functional elements from run-on sequencing data remains an open problem because the signal is noisy and challenging to model. Here we investigated computational approaches that convert run-on and sequencing data into annotations representing transcription units, including genes and non-coding RNAs. We developed a convolutional neural network, called convolutional discovery of gene anatomy using PRO-seq (CGAP), trained to identify different anatomical features of a transcription unit, which were then stitched together into transcript annotations using a hidden Markov model (HMM). Comparison with existing methods showed a significant performance improvement using our novel CGAP-HMM approach. We developed a voting system that ensembles the top three annotation strategies, resulting in large and significant improvements in transcription unit annotation accuracy over the best performing individual method. Finally, we also explore a conditional generative adversarial network (cGAN) as a possible alternative approach to transcription unit annotation. Collectively our work provides novel tools for de novo transcription unit annotation from run-on and sequencing data that are accurate enough to be useful in many applications.Author summary: Understanding how transcriptional elements (e.g., genes) are organized and expressed is fundamental to biology and medicine. Our DNA contains thousands of transcriptional elements, but pinpointing exactly where each begins and ends (a process called genome annotation) remains technically challenging, especially for newly sequenced organisms. One powerful experimental approach, called precision nuclear run-on and sequencing (PRO-seq), captures RNA polymerase (the molecular machine that reads DNA to produce RNA) as it works across the genome, generating a characteristic “signal fingerprint” at each active transcriptional element.
Suggested Citation
Paul R Munn & Jay Chia & Charles G Danko, 2026.
"Accurate de novo transcription unit annotation from run-on and sequencing data,"
PLOS Computational Biology, Public Library of Science, vol. 22(8), pages 1-30, August.
Handle:
RePEc:plo:pcbi00:1014559
DOI: 10.1371/journal.pcbi.1014559
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014559. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.