Skip to content

Author

Henry De Courcy Thompson

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#edge computing Open access Sep 2026

Approximate Equilibrium Play in 4-Max Shove-or-Fold Poker: A Delayed-Linear Weighted-Averaging MCCFR Study at 8BB Effective Stacks

This technical note presents validated shove and call ranges for 4-player all-in-or-fold poker at 8 big-blind effective stacks. The strategy was produced using delayed-linear weighted-averaging Monte Carlo Counterfactual Regret Minimization (MCCFR), trained from iteration zero to 4,000,000 across two independent seeds and merged through the FullHouse production adapter. The resulting Production v2 policy passed a four-stage adoption protocol incorporating structural validation, NashConv evaluation, whole-game comparison against the previous production policy, and a frozen stress-cell regression. The latter returned a mixed result and is retained as adverse evidence. The published ranges should therefore be interpreted as a validated finite-compute estimate of equilibrium play rather than a formal proof of Nash equilibrium. Mixed-strategy frequencies displayed in the PDF are rounded to the nearest 25% for readability; the underlying policy retains exact frequencies. Feedback, corrections, methodological criticism and edge-case reports are welcome.

Henry De Courcy Thompson · 0 citations
#edge computing Open access Sep 2026

Approximate Equilibrium Play in 4-Max Shove-or-Fold Poker: A Delayed-Linear Weighted-Averaging MCCFR Study at 8BB Effective Stacks

This technical note presents validated shove and call ranges for 4-player all-in-or-fold poker at 8 big-blind effective stacks. The strategy was produced using delayed-linear weighted-averaging Monte Carlo Counterfactual Regret Minimization (MCCFR), trained from iteration zero to 4,000,000 across two independent seeds and merged through the FullHouse production adapter. The resulting Production v2 policy passed a four-stage adoption protocol incorporating structural validation, NashConv evaluation, whole-game comparison against the previous production policy, and a frozen stress-cell regression. The latter returned a mixed result and is retained as adverse evidence. The published ranges should therefore be interpreted as a validated finite-compute estimate of equilibrium play rather than a formal proof of Nash equilibrium. Mixed-strategy frequencies displayed in the PDF are rounded to the nearest 25% for readability; the underlying policy retains exact frequencies. Feedback, corrections, methodological criticism and edge-case reports are welcome.

Henry De Courcy Thompson · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.