IDEAS home Printed from https://ideas.repec.org/a/plo/pcbi00/1014723.html

Deterministic dynamics of distributional multi-agent reinforcement learning

Author

Listed:
  • Clémence Bergerot
  • Pawel Romanczuk
  • Wolfram Barfuss

Abstract

Understanding how cognition shapes behavior across contexts remains a fundamental challenge for many disciplines. In particular, for the optimism heuristic–i.e., the tendency to overweight positive (relative to negative) information–knowledge remains fragmented, with models developed in specific domains in isolation. Here, we present a unifying computational framework by deriving the deterministic dynamics of distributional multi-agent reinforcement learning. Our approach discretizes return distributions through a finite set of neurons, consistent with recent empirical findings on distributional coding in the brain. We validate our framework by reproducing established results across three iconic domains spanning individual bandit choice under resource variability, social coordination, and risky choice. Beyond validation, we uncover novel interactions among optimism, return discretization, and temporal discounting. Specifically, we identify conditions under which return discretization generates choice hysteresis and, in extreme parameter regimes, inescapable perseveration. We further reveal “individual dilemmas”: circumstances where agents gravitate toward suboptimal yet stable strategies, offering a mechanistic explanation for incoherent choice patterns. Our framework bridges neuroscience, psychology, and collective behavior, enabling empirically testable hypotheses about how cognitive biases propagate from individual cognition to social outcomes in complex environments.Author summary: How do cognitive heuristics shape individual decisions and the collective outcomes that emerge when many agents interact? This question connects psychology, neuroscience, and artificial intelligence, yet computational models have developed separately across these fields. Here we present DDRL, a framework that combines distributional reinforcement learning, in which agents learn the full spread of possible outcomes rather than their average, with deterministic learning dynamics that yield mathematically tractable trajectories in strategy space. DDRL is grounded in neuroscientific evidence: dopaminergic neurons encode outcome distributions rather than expected values alone. We illustrate DDRL with an optimism heuristic across three domains: bandit choice under resource variability, social coordination, and intertemporal risky choice. In each case, DDRL reproduces established results. Beyond replication, DDRL makes two novel predictions. Return discretization generates bistable regimes, in which agents settle into qualitatively different behaviors depending on their history. Under certain combinations of optimism and temporal discounting, agents can become locked into an individual dilemma, in which the better strategy is no longer a stable attractor. Together, these results provide a mechanistic route from neural reward encoding to emergent collective behavior, connecting individual cognitive heuristics to social outcomes in a biologically grounded and analytically tractable framework.

Suggested Citation

  • Clémence Bergerot & Pawel Romanczuk & Wolfram Barfuss, 2026. "Deterministic dynamics of distributional multi-agent reinforcement learning," PLOS Computational Biology, Public Library of Science, vol. 22(9), pages 1-24, September.
  • Handle: RePEc:plo:pcbi00:1014723
    DOI: 10.1371/journal.pcbi.1014723
    as

    Download full text from publisher

    File URL: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014723
    Download Restriction: no

    File URL: https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1014723&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pcbi.1014723?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014723. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.