Skip to content
#edge computing Open access

SESTINA v1.0: Exact Inference, Robust Evaluation, and Adversarial Testing in Six-Player Canadian Fish

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics

Abstract

Six-player Canadian Fish is a decentralised imperfect-information team game in which every action is public, so the hidden state reduces to the initial deal and the posterior over that deal can be computed exactly. We develop and evaluate SESTINA v1.0, the strongest configuration produced in this project’s FishBot lineage. It combines an approximate Sinkhorn fit to that posterior, started from a fitted policy prior, with a linear ask and declaration policy, a public-history tie-breaking rule that preserves common knowledge among teammates, a half-suit contestation weighting, a deduction-state stall detector in place of an event-count termination rule, and a guarded determinized test-time search. Evaluation follows a protocol registered in advance and run on sealed holdout material, using duplicate deal blocks, deal-clustered bootstrap confidence intervals, replication across two disjoint deal banks as an advance-specified criterion, calibrated detection floors, and mechanical side-channel controls. Against F-cheap, the cheapest configuration genuinely on the v0.6 frontier and the registered comparison target, SESTINA v1.0 achieves a +3.33 percentage-point win-rate edge (95% CI [+2.88, +3.78]) over 48,000 sealed games, with the sign replicating on both banks. The advantage persists under cross-play between independently trained runs and eight rule dialects. Under partner substitution against a v0.5 opponent 6 of 7 changed-partner rows stay positive, but the worst is −0.19 [−1.04, +0.67], which that battery does not resolve against the registered −1.00 collapse threshold. Separately, none of eight independently constructed adversarial searches found a positive edge at the tested budgets. SESTINA v1.0 does not, however, measurably outperform a composite configuration assembled earlier in the same programme (+0.15 pp, 95% CI [−0.29, +0.59]), indicating that the architecture work which followed added no measurable strength. Over a shared 31-member opponent panel SESTINA v1.0’s worst cell is −0.04 pp [−1.41, +1.33], which does not replicate in sign, and it is 3rd of four on minimax regret. Four candidate mechanisms failed to produce measurable improvement at this resolution. At the calibrated resolution of this evaluation the tested policy class appears locally flat, and we state what evidence a near-optimality claim would require.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.