1. Auer, P., Cesa-Bianchi, N., & Fischer, P. (2002). Finite-time analysis of the multiarmed Bandit problem. Machine Learning, 47(2/3), 235–256.
2. Axelrod, R., & Hamilton, W. D. (1981). The evolution of cooperation. Science, 211(27), 1390–1396.
3. Babes, M., Munoz de Cote, E., & Littman, M. L. (2008). Social reward shaping in the prisoner’s dilemma. In Proceedings of the 7th International Conference on Autonomous Agents and Multiagent Systems, (pp. 1389–1392). Estoril: International Foundation for Autonomous Agents and Multiagent Systems.
4. Banerjee, B., & Peng, J. (2005). Efficient learning of multi-step best response. In Proceedings of the 4th International Conference on Autonomous Agents and Multiagent Systems, (pp. 60–66). Utretch: ACM.
5. Bard, N., Johanson, M., Burch, N., & Bowling, M. (2013). Online implicit agent modelling. In Proceedings of the 12th International Conference on Autonomous Agents and Multiagent Systems, (pp. 255–262). Saint Paul, MN: International Foundation for Autonomous Agents and Multiagent Systems.