Author
Listed:
- William L Tong
- Venkatesh N Murthy
- Gautam Reddy
Abstract
Dogs and laboratory mice are commonly trained to perform complex tasks by guiding them through a curriculum of simpler tasks (‘shaping’). What are the principles behind effective shaping strategies? Here, we propose a teacher-student framework for shaping behavior, where an autonomous teacher agent decides its student’s task based on the student’s transcript of successes and failures on previously assigned tasks. Using algorithms for Monte Carlo planning under uncertainty, we show that near-optimal shaping algorithms achieve a careful balance between reinforcement and extinction. Near-optimal algorithms track learning rate to adaptively alternate between simpler and harder tasks. Based on this intuition, we derive an adaptive shaping heuristic with minimal parameters, which we show is near-optimal on a sequence learning task and robustly trains deep reinforcement learning agents on navigation tasks that involve sparse, delayed rewards. Extensions to continuous curricula are explored. Our work provides a starting point towards a general computational framework for shaping behavior that applies to both animals and artificial agents.Author summary: Animals are commonly trained by ‘shaping’ their behavior using a sequence of simpler tasks towards a complex behavior. Numerous schools of thought have proposed heuristics for shaping based on qualitative principles of reinforcement learning. We introduce a general computational framework for shaping behavior, paying special attention to the constraints faced when training animals. Using machine learning algorithms for planning under uncertainty, we explain why simple strategies fail, provide a normative foundation for existing heuristics, and propose new adaptive algorithms for designing curricula.
Suggested Citation
William L Tong & Venkatesh N Murthy & Gautam Reddy, 2025.
"Adaptive algorithms for shaping behavior,"
PLOS Computational Biology, Public Library of Science, vol. 21(9), pages 1-15, September.
Handle:
RePEc:plo:pcbi00:1013454
DOI: 10.1371/journal.pcbi.1013454
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1013454. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.