AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement Learning-Reference-Cited by-同舟云学术

AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement Learning

Published:2022-06-28 Issue:4 Volume:36 Page:4689-4697
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Zhao Enmin,Yan Renye,Li Jinqiu,Li Kai,Xing Junliang

Abstract

Heads-up no-limit Texas hold’em (HUNL) is the quintessential game with imperfect information. Representative priorworks like DeepStack and Libratus heavily rely on counter-factual regret minimization (CFR) and its variants to tackleHUNL. However, the prohibitive computation cost of CFRiteration makes it difficult for subsequent researchers to learnthe CFR model in HUNL and apply it in other practical applications. In this work, we present AlphaHoldem, a high-performance and lightweight HUNL AI obtained with an end-to-end self-play reinforcement learning framework. The proposed framework adopts a pseudo-siamese architecture to directly learn from the input state information to the output actions by competing the learned model with its different historical versions. The main technical contributions include anovel state representation of card and betting information, amultitask self-play training loss function, and a new modelevaluation and selection metric to generate the final model.In a study involving 100,000 hands of poker, AlphaHoldemdefeats Slumbot and DeepStack using only one PC with threedays training. At the same time, AlphaHoldem only takes 2.9milliseconds for each decision-making using only a singleGPU, more than 1,000 times faster than DeepStack. We release the history data among among AlphaHoldem, Slumbot,and top human professionals in the author’s GitHub repository to facilitate further studies in this direction.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 14 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Capitalizing on the Opponent's Uncertainty in Reconnaissance Blind Chess;2024 IEEE Congress on Evolutionary Computation (CEC);2024-06-30

2. Transformer in reinforcement learning for decision-making: a survey;Frontiers of Information Technology & Electronic Engineering;2024-06

3. Clicked:Curriculum Learning Connects Knowledge Distillation for Four-Player No-Limit Texas Hold’em Poker;2024 36th Chinese Control and Decision Conference (CCDC);2024-05-25

4. Survey of Opponent Modeling: from Game AI to Combat Deduction;2024 36th Chinese Control and Decision Conference (CCDC);2024-05-25

5. An improved deep Q-Network algorithm for the prediction of non-competitive bidding in Bridge Game;Proceedings of the 2024 5th International Conference on Computing, Networks and Internet of Things;2024-05-24