Reinforcement Learning Dynamics in Social Dilemmas
In this paper we replicate and advance Macy and Flache's (2002; Proc. Natl. Acad. Sci. USA, 99, 7229â€“7236) work on the dynamics of reinforcement learning in 2×2 (2-player 2-strategy) social dilemmas. In particular, we provide further insight into the solution concepts that they describe, illustrate some recent analytical results on the dynamics of their model, and discuss the robustness of such results to occasional mistakes made by players in choosing their actions (i.e. trembling hands). It is shown here that the dynamics of their model are strongly dependent on the speed at which players learn. With high learning rates the system quickly reaches its asymptotic behaviour; on the other hand, when learning rates are low, two distinctively different transient regimes can be clearly observed. It is shown that the inclusion of small quantities of randomness in players' decisions can change the dynamics of the model dramatically.
Volume (Year): 11 (2008)
Issue (Month): 2 ()
|Contact details of provider:|| |
References listed on IDEAS
Please report citation or reference errors to , or , if you are the registered author of the cited work, log in to your RePEc Author Service profile, click on "citations" and make appropriate adjustments.:
- Andreas Flache & Michael W. Macy, 2002. "Stochastic Collusion and the Power Law of Learning," Journal of Conflict Resolution, Peace Science Society (International), vol. 46(5), pages 629-653, October.
- Barry Sopher & Dilip Mookherjee, 2000.
"Learning and Decision Costs in Experimental Constant Sum Games,"
Departmental Working Papers
199625, Rutgers University, Department of Economics.
- Mookherjee, Dilip & Sopher, Barry, 1997. "Learning and Decision Costs in Experimental Constant Sum Games," Games and Economic Behavior, Elsevier, vol. 19(1), pages 97-132, April.
- Barry Sopher & Dilip Mookherjee, 1997. "Learning and Decision Costs in Experimental Constant Sum Games," Departmental Working Papers 199527, Rutgers University, Department of Economics.
- Binmore, K. & Samuelson, L., 1993. "An Economist's Perspective on the Evolution of Norms," Working papers 9323, Wisconsin Madison - Social Systems.
- John G. Cross, 1973. "A Stochastic Learning Model of Economic Behavior," The Quarterly Journal of Economics, Oxford University Press, vol. 87(2), pages 239-266.
- Glenn Ellison, 2000. "Basins of Attraction, Long-Run Stochastic Stability, and the Speed of Step-by-Step Evolution," Review of Economic Studies, Oxford University Press, vol. 67(1), pages 17-45.
- J. Gary Polhill & Luis R. Izquierdo, 2005. "Lessons Learned from Converting the Artificial Stock Market to Interval Arithmetic," Journal of Artificial Societies and Social Simulation, Journal of Artificial Societies and Social Simulation, vol. 8(2), pages 1-2.
- Margaret Edwards & Sylvie Huet & FranÃ§ois Goreaud & Guillaume Deffuant, 2003. "Comparing an Individual-Based Model of Behaviour Diffusion with Its Mean Field Aggregate Approximation," Journal of Artificial Societies and Social Simulation, Journal of Artificial Societies and Social Simulation, vol. 6(4), pages 1-9.
- Tilman Börgers & Rajiv Sarin, "undated".
"Learning Through Reinforcement and Replicator Dynamics,"
ELSE working papers
051, ESRC Centre on Economics Learning and Social Evolution.
- Borgers, Tilman & Sarin, Rajiv, 1997. "Learning Through Reinforcement and Replicator Dynamics," Journal of Economic Theory, Elsevier, vol. 77(1), pages 1-14, November.
- T. Borgers & R. Sarin, 2010. "Learning Through Reinforcement and Replicator Dynamics," Levine's Working Paper Archive 380, David K. Levine.
- Sylvie Huet & Margaret Edwards & Guillaume Deffuant, 2007. "Taking into Account the Variations of Neighbourhood Sizes in the Mean-Field Approximation of the Threshold Model on a Random Network," Journal of Artificial Societies and Social Simulation, Journal of Artificial Societies and Social Simulation, vol. 10(1), pages 1-10.
- Erev, Ido & Bereby-Meyer, Yoella & Roth, Alvin E., 1999. "The effect of adding a constant to all payoffs: experimental investigation, and implications for reinforcement learning models," Journal of Economic Behavior & Organization, Elsevier, vol. 39(1), pages 111-128, May.
- Mookherjee Dilip & Sopher Barry, 1994. "Learning Behavior in an Experimental Matching Pennies Game," Games and Economic Behavior, Elsevier, vol. 7(1), pages 62-91, July.
- Roth, Alvin E. & Erev, Ido, 1995. "Learning in extensive-form games: Experimental data and simple dynamic models in the intermediate term," Games and Economic Behavior, Elsevier, vol. 8(1), pages 164-212.
- Bendor Jonathan & Mookherjee Dilip & Ray Debraj, 2001. "Reinforcement Learning in Repeated Interaction Games," The B.E. Journal of Theoretical Economics, De Gruyter, vol. 1(1), pages 1-44, March.
- Karandikar, Rajeeva & Mookherjee, Dilip & Ray, Debraj & Vega-Redondo, Fernando, 1998.
"Evolving Aspirations and Cooperation,"
Journal of Economic Theory,
Elsevier, vol. 80(2), pages 292-331, June.
- Luis R. Izquierdo & J. Gary Polhill, 2006. "Is Your Model Susceptible to Floating-Point Errors?," Journal of Artificial Societies and Social Simulation, Journal of Artificial Societies and Social Simulation, vol. 9(4), pages 1-4.
- Fernando Vega-Redondo & Frédéric Palomino, 1999.
"Convergence of aspirations and (partial) cooperation in the prisoner's dilemma,"
International Journal of Game Theory,
Springer;Game Theory Society, vol. 28(4), pages 465-488.
- Palomino, F. & Vega, F., 1996. "Convergence of Aspirations and (Partial) Cooperation in the Prisoners's Dilemma," UFAE and IAE Working Papers 345.96, Unitat de Fonaments de l'Anàlisi Econòmica (UAB) and Institut d'Anàlisi Econòmica (CSIC).
- Fernando Vega Redondo & Frédéric Palomino, 1996. "Convergence of aspirations and (partial) cooperation in the Prisoner's Dilemma," Working Papers. Serie AD 1996-20, Instituto Valenciano de Investigaciones Económicas, S.A. (Ivie).
- Yan Chen & Fang-Fang Tang, 1998. "Learning and Incentive-Compatible Mechanisms for Public Goods Provision: An Experimental Study," Journal of Political Economy, University of Chicago Press, vol. 106(3), pages 633-662, June.
- Youngse Kim, 1999. "Satisficing and optimality in 2þ2 common interest games," Economic Theory, Springer;Society for the Advancement of Economic Theory (SAET), vol. 13(2), pages 365-375.
- JosÃ© Manuel GalÃ¡n & Luis R. Izquierdo, 2005. "Appearances Can Be Deceiving: Lessons Learned Re-Implementing Axelrod's 'Evolutionary Approach to Norms'," Journal of Artificial Societies and Social Simulation, Journal of Artificial Societies and Social Simulation, vol. 8(3), pages 1-2.
- Izquierdo, Luis R. & Izquierdo, Segismundo S. & Gotts, Nicholas M. & Polhill, J. Gary, 2007. "Transient and asymptotic dynamics of reinforcement learning in games," Games and Economic Behavior, Elsevier, vol. 61(2), pages 259-276, November.
When requesting a correction, please mention this item's handle: RePEc:jas:jasssj:2007-11-2. See general information about how to correct material in RePEc.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: (Flaminio Squazzoni)
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
If references are entirely missing, you can add them using this form.
If the full references list an item that is present in RePEc, but the system did not link to it, you can help with this form.
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your profile, as there may be some citations waiting for confirmation.
Please note that corrections may take a couple of weeks to filter through the various RePEc services.