Author
Listed:
- Junfan Chen
- Fabian Schmidt
- Ricardo Henao
Abstract
Single-cell transcriptomic data provide critical insights into cellular states and disease mechanisms, and foundation models have recently emerged as powerful tools for learning gene–gene relationships from these data. However, current approaches often overlook key challenges, including the mismatch between model design and the rank-ordered structure of gene expression profiles, as well as the unclear benefits of large-scale pretraining for biological applications. Here, we present GFCAB, a modified modeling framework designed to better capture the structural properties of ranked single-cell transcriptomic data. GFCAB incorporates a cumulative assignment mechanism to suppress repeated gene predictions and a similarity-based regularization strategy to promote diversity in model outputs. Across multiple evaluation settings, including pretraining behavior, biologically relevant classification tasks, and cross-dataset analyzes, GFCAB consistently reduces redundancy and enhances the recovery of low-frequency genes with known functional and disease relevance while maintaining or improving predictive accuracy. In downstream applications, including classification and zero-shot batch effect correction, the model achieves competitive or improved performance compared to existing approaches. We further show that indiscriminately increasing the pretraining data scale does not uniformly improve performance. Instead, models trained on substantially smaller datasets can match or exceed the performance of larger models and often demonstrate improved generalization across datasets. Together, these findings highlight the importance of aligning model design with the intrinsic structure of biological data and suggest that architectural innovation can reduce reliance on large-scale training data. GFCAB provides a framework for developing more efficient and biologically informative models for single-cell analysis, with potential applications in disease characterization and precision biology.Author summary: Single-cell analysis helps researchers understand how genes work together inside individual cells, and recent artificial intelligence models have shown strong potential for uncovering these patterns. However, many existing approaches do not fully account for how this data is structured, and often assume that using more training data will always improve performance. In this study, we introduce GFCAB, a model designed to better match the way single-cell data are organized. By reducing repeated predictions and encouraging the model to consider a wider range of genes, GFCAB produces more diverse and biologically meaningful results, including the identification of less common but important gene signals. We also show that simply increasing the amount of training data does not always lead to better outcomes. Instead, well-designed models trained on smaller datasets can perform just as well and often generalize better to new data. These findings highlight the importance of model design in developing efficient and reliable tools for single-cell research.
Suggested Citation
Junfan Chen & Fabian Schmidt & Ricardo Henao, 2026.
"Assessing scale and predictive diversity in models for single-cell transcriptomics based on Geneformer,"
PLOS Computational Biology, Public Library of Science, vol. 22(7), pages 1-21, July.
Handle:
RePEc:plo:pcbi00:1013701
DOI: 10.1371/journal.pcbi.1013701
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1013701. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.