Abstract
Abstract
English speakers use probabilistic phrases such as likely to communicate information about the probability or likelihood of events. Communication is successful to the extent that the listener grasps what the speaker means to convey and, if communication is successful, individuals can potentially coordinate their actions based on shared knowledge about uncertainty. We first assessed human ability to estimate the probability and the ambiguity (imprecision) of twenty-three probabilistic phrases in a coordination game in two different contexts, investment advice and medical advice. We then had GPT-4 (OpenAI), a Large Language Model, complete the same tasks as the human participants. We found that GPT-4’s estimates of probability both in the investment and Medical contexts were as close or closer to that of the human participants as the human participants’ estimates were to one another. However, further analyses of residuals disclosed small but significant differences between human and GPT-4 performance. In particular, human probability estimates were compressed relative to those of GPT-4. Estimates of probability for both the human participants and GPT-4 were little affected by context. We propose that evaluation methods based on coordination games provide a systematic way to assess what GPT-4 and similar programs can and cannot do.
Publisher
Research Square Platform LLC
Reference51 articles.
1. Austin, J. L. (1955), How to do things with words. Oxford: Oxford University Press.
2. Benz, A., Ebert, C., Jäger, G. & Van Rooij, R. [Eds.] (2011), Language, games and evolution; Trends in current research on language and game theory. Berlin Heidelberg: Springer.
3. Benz, A., Jäger, G. & Van Rooij, R. [Eds.] (2014),Game theory and pragmatics. Springer. Berlin Heidelberg.
4. Benz, A. & Stevens, J. (2018), Game-theoretic approaches to pragmatics. Annual Review of Linguistics, 4, 173 − 91.
5. Beyth-Marom, R (1982), How probable is probable? A numerical translation of verbal probability expressions. Journal of Forecasting. 1, 257–269.