Mixed-Strategy Decision Tree (MDT) is proposed, which articulates the silent optimality of the equilibrium into sparse strategic rules that both humans and LLMs could understand and extends the input to arbitrarily new states and continuations.
Abstract
Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. In fact, human play is often guided by intuition and heuristics and can deviate substantially from game equilibrium. This discrepancy is amplified in games with mixed-strategy equilibria, where human data is heavily biased toward pure strategies. Consequently, conditioning LLMs on this data yields weak game strategies. To grant LLMs the reasoning capacity in games, in this work, we study how to elicit equilibrium play using solver output. We propose Mixed-Strategy Decision Tree (MDT), which articulates the silent optimality of the equilibrium into sparse strategic rules that both humans and LLMs could understand. Using solver output rather than human annotation allows us to extend the input to arbitrarily new states and continuations. We instantiate this study on No-Limit Texas Hold'em by querying a solver oracle for over \textbf{250 million mixed-strategy decisions}; MDT together with other techniques \textbf{reduces the $\ell_1$ distance to the equilibrium by $52.6\%$} across $8$ different LLM configurations. A Route-only ablation tests the incremental contribution of the shadow-based contrast, while complete River-endgame and Liar's Dice experiments evaluate strategic fidelity and portability beyond the original NLH communication setting.
This work formalises a necessary level-K distinguishability condition for strategic depth inference and builds a suite of novel game structures that meet this standard, and finds that model models maintain accurate strategic depth under recursive reasoning, with strong internal consistency between stated reasoning and actions at every level.
CAST (Credit Assignment from Solver Teachers), which converts value changes in a game solver's state value into solver advantages and injects them into RLVR as turn-level signals and achieves the highest average zero-shot performance on ALFWorld and WebShop.
Yu Wang, Yi-Kai Zhang, Wentao Shi et al.· 0 citations
This work identifies narrow-support imitation as a source of policy collapse in LLM decision-making and suggests that preserving action support during SFT is important for maintaining exploratory behavior.
Junyi Sha, Renfei Tan, David Simchi-Levi· arXiv.org· 0 citations
This article introduces a framework for designing and running simulated experiments with LLM‐powered agents and applies the framework to the exploration–exploitation dilemma and shows that LLM‐based experiments reproduce patterns observed among human participants.
The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential branching.
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al.· 1 citation
An agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment is introduced, and it is shown that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.
Jakub Rada, Viliam Lisý AI Center, Department of Rehabilitation Science et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.