IDEAS home Printed from https://ideas.repec.org/a/plo/pcbi00/1014570.html

NLCD: A method to discover nonlinear causal relations among genes

Author

Listed:
  • Aravind Easwar
  • Manikandan Narayanan

Abstract

Distinguishing correlation from causation is a fundamental challenge in many scientific fields, including biology, especially when interventions like randomized controlled trials are infeasible and only observational data are available. Methods based on statistical tests of conditional independence within the Mendelian Randomization framework can detect causality between two observed variables that are each associated with a third instrumental variable. However, these methods for detecting causal relationships between traits (e.g., two gene expression or clinical traits associated with a genetic variant, all observed in the same population) often assume a linear relationship, thereby hindering the discovery of causal gene networks from genomics data.We have developed NLCD, a method for NonLinear Causal Discovery from genomics data based on nonlinear regression modeling and conditional feature importance scoring. NLCD uses these techniques to extend the statistical tests in an existing linear causal discovery method called the Causal Inference Test (CIT). We benchmarked NLCD against current state-of-the-art methods: CIT, Findr, and MRPC. On simulated datasets, NLCD performs comparably to most methods in detecting linear relations (Average AUPRC (Area Under the Precision-Recall Curve) of NLCD = 0.94, CIT = 0.94, Findr = 0.94, and MRPC = 0.99), and outperforms them in detecting nonlinear (sine and sawtooth type) relations between two genes (Average AUPRC of NLCD = 0.76, CIT = 0.60, Findr = 0.56, and MRPC = 0.73). When tested on a nonlinear subset of a yeast genomic dataset to recover known causal relations involving transcription factors, NLCD and CIT performed comparable to each other and slightly better than Findr and MRPC (Average AUPRC of NLCD = 0.82, CIT = 0.81, Findr = 0.71, and MRPC = 0.54). On application to a human genomic dataset, NLCD revealed active causal gene pairs (IRF1 → PSME1 and HLA-C → HLA-T) in the muscle tissue, and clarified the promises and challenges in discovering causal gene networks in tissues under in vivo human settings.Author summary: A central goal in biology is to discover causal mechanisms. But establishing causality between two biological factors (such as two genes) is difficult when interventions are expensive, impractical, or impossible as in human populations. Researchers therefore rely on observational genetic and other genomic data of hundreds of individuals to infer causal relationships between genes or other factors. Existing computational methods addressing this problem often assume that biological effects follow simple linear patterns. In reality, biological systems can exhibit more complex behaviors, which can cause important causal relationships to be overlooked. To address this limitation, we have developed a method that can detect causal relationships even when they are nonlinear. Our approach builds on existing causal discovery frameworks while using machine learning techniques to capture more complex patterns in the data. In extensive tests on simulated datasets, our method performs as well as current approaches for linear relationships and substantially better when relationships are nonlinear. When applied on yeast and human genomic datasets, our method recovered known causal interactions and identified biologically relevant gene pairs. The causal relations predicted by our method can serve as hypotheses to guide future experimental studies.

Suggested Citation

  • Aravind Easwar & Manikandan Narayanan, 2026. "NLCD: A method to discover nonlinear causal relations among genes," PLOS Computational Biology, Public Library of Science, vol. 22(7), pages 1-30, July.
  • Handle: RePEc:plo:pcbi00:1014570
    DOI: 10.1371/journal.pcbi.1014570
    as

    Download full text from publisher

    File URL: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014570
    Download Restriction: no

    File URL: https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1014570&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pcbi.1014570?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014570. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.