PANES: Policy-guided Asymmetric Nash Equilibrium Search
Trick-taking card games with mandatory bidding confront reinforcement learning agents with a distinctive two-phase problem: each player must commit to a numeric bid before any cards are played, and whether that bid turns out to be correct hinges on adversarial interactions unfolding over many subsequent tricks. Judgeme...