Author
Abstract
Machine learning (ML) methods for proteins and RNAs rely on multiple sequence alignments (MSAs) and related datasets such as experimental mutagenesis libraries, yet the amount of usable information they contain remains unclear. Here, a spectral measure of information is recast into an interpretable quantity for MSAs, denoted Leff, defined as the number of fully independent alignment positions that reproduce the observed sequence diversity. Applied to RNA MSAs, this measure shows that evolutionary constraints nearly halve diversity relative to the secondary structure alone, quantifying functional and phylogenetic restrictions beyond base pairing. The same analysis indicates even lower effective diversity in proteins, reflecting tighter packing and coevolutionary constraints. Leff further correlates with protein structure prediction accuracy, anticipating cases with insufficient evolutionary signal. When applied to experimentally and computationally generated libraries, it measures both produced diversity and cross-library overlap, quantifying novelty rather than redundant sampling. Together, these results establish Leff as an operational tool to estimate effective information in MSAs, anticipate modeling difficulties, and guide protein and RNA design.Author summary: Machine learning has transformed biology, predicting protein structures, uncovering evolutionary rules, and designing new RNA and protein sequences. Almost every such method learns from large collections of related sequences, and the field largely assumes that more data means better models. But more is not always richer. A collection of thousands of sequences may hold far fewer independent evolutionary signals, because so many entries are near-copies shaped by shared ancestry or by designs that scarcely depart from a single template. We rarely know how much real information a dataset carries, let alone how to measure it. Here I introduce the effective length, a simple measure of the genuinely independent information in a sequence collection. It reveals that natural RNA and protein families are far more constrained than their size suggests, anticipates how reliably their structures can be predicted, and distinguishes new sequence libraries that add real information from those that merely repeat what we already know.
Suggested Citation
Vaitea Opuu, 2026.
"A spectral framework for measuring diversity in multiple sequence alignments,"
PLOS Computational Biology, Public Library of Science, vol. 22(9), pages 1-17, September.
Handle:
RePEc:plo:pcbi00:1014778
DOI: 10.1371/journal.pcbi.1014778
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014778. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.