Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while na...
The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Because ICM reads only stack sizes, it omits action order, blind obligations, and seat rotation, and it does not price the elimination pressure a big stack puts on the sho...
GPU-CFR is proposed, a compiler and runtime built on observation that for a fixed game, everything about a CFR iteration except the numerical values is known before the first iteration runs, and beats every CPU and GPU baseline on the mid-to-large games of the suite without changing the update rule.
The Abstraction Agent is proposed, a zero-shot pipeline that uses a large language model (LLM) to discover continuous strategic features from a natural-language game description, score private states on these features, and cluster them into abstraction buckets, without any game-specific evaluator, training data, or gam...
CS-RNR is introduced, the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself.
Bo-Ning Li, Long-Bo Huang· arXiv.org· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.