Probability Matching and Reinforcement Learning
Probability matching occurs when an action is chosen with a frequency equivalent to the probability of that action being the best choice. This sub-optimal behavior has been reported repeatedly by psychologist and experimental economist. We provide an evolutionary foundation for this phenomenon by showing that learning by reinforcement can lead to probability matching and, if learning occurs suffciently slowly, probability matching does not only occur in choice frequencies but also in choice probabilities. Our results are completed by proving that there exists no quasi-linear reinforcement learning specification such that behavior is optimal for all environments where counterfactuals are observed.
|Date of creation:||Mar 2011|
|Date of revision:|
|Contact details of provider:|| Postal: Department of Economics University of Leicester, University Road. Leicester. LE1 7RH. UK|
Phone: +44 (0)116 252 2887
Fax: +44 (0)116 252 2908
Web page: http://www2.le.ac.uk/departments/economics
More information through EDIRC
|Order Information:|| Web: http://www2.le.ac.uk/departments/economics/research/discussion-papers Email: |
Please report citation or reference errors to , or , if you are the registered author of the cited work, log in to your RePEc Author Service profile, click on "citations" and make appropriate adjustments.:
- Tilman Börgers & Rajiv Sarin, .
"Naive Reinforcement Learning With Endogenous Aspiration,"
ELSE working papers
037, ESRC Centre on Economics Learning and Social Evolution.
- Borgers, Tilman & Sarin, Rajiv, 2000. "Naive Reinforcement Learning with Endogenous Aspirations," International Economic Review, Department of Economics, University of Pennsylvania and Osaka University Institute of Social and Economic Research Association, vol. 41(4), pages 921-50, November.
- T. Borgers & R. Sarin, 2010. "Naïve Reinforcement Learning With Endogenous Aspirations," Levine's Working Paper Archive 381, David K. Levine.
- Roth, Alvin E. & Erev, Ido, 1995. "Learning in extensive-form games: Experimental data and simple dynamic models in the intermediate term," Games and Economic Behavior, Elsevier, vol. 8(1), pages 164-212.
- Javier Rivas, 2008. "Learning within a Markovian Environment," Economics Working Papers ECO2008/13, European University Institute.
- Erev, Ido & Roth, Alvin E, 1998. "Predicting How People Play Games: Reinforcement Learning in Experimental Games with Unique, Mixed Strategy Equilibria," American Economic Review, American Economic Association, vol. 88(4), pages 848-81, September.
- Kosfeld, Michael & Droste, Edward & Voorneveld, Mark, 2002.
"A myopic adjustment process leading to best-reply matching,"
Games and Economic Behavior,
Elsevier, vol. 40(2), pages 270-298, August.
- Droste, E.J.R. & Kosfeld, M. & Voorneveld, M., 1998. "A Myopic Adjustment Process Leading to Best-Reply Matching," Discussion Paper 1998-111, Tilburg University, Center for Economic Research.
- Colin Camerer & Teck-Hua Ho, 1999. "Experience-weighted Attraction Learning in Normal Form Games," Econometrica, Econometric Society, vol. 67(4), pages 827-874, July.
- Samuelson Larry, 1994. "Stochastic Stability in Games with Alternative Best Replies," Journal of Economic Theory, Elsevier, vol. 64(1), pages 35-65, October.
- Rubinstein, Ariel, 2002. "Irrational diversification in multiple decision problems," European Economic Review, Elsevier, vol. 46(8), pages 1369-1378, September.
When requesting a correction, please mention this item's handle: RePEc:lec:leecon:11/20. See general information about how to correct material in RePEc.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: (Mrs. Alexandra Mazzuoccolo)
If references are entirely missing, you can add them using this form.