This work tests a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game.
Abstract
Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\geq 0.5$) about the hidden piece's location were correct in 1 of 62 cases. The calibration deficit is concentrated almost entirely in these events: 99.3% of it in the original batch, 98.7% in the replication. The same pattern, in weaker form, orders consistently (point estimates only; most pairwise gaps are not statistically distinguishable at this sample size) across four further model configurations spanning a second provider -- reported as scope for the finding, not as evidence that capability predicts calibration: a same-model comparison at a fixed external leaderboard score shows a deliberation-budget change alone moves the metric by nearly as much as a large cross-model gap. In a separate seat, conventional evaluation axes -- legality, cost, latency, completion rate -- can dissociate entirely from belief quality, with the configuration winning on every conventional axis producing the worst belief quality tested. A model exhibiting this pattern can still win the game its belief was about, which is why outcome-only evaluation would not detect it.
An auditable framework is built that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change.
Yuanpeng Gao, Jiangyi Yang, Yao Zhao et al.· arXiv.org· 0 citations
MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind, is presented, finding that stated confidence is poorly calibrated, agents overestimate how often they are suspected by a factor of 1.5, and single-vote counterfactual replays rarely change game outcomes.
Early deception poses a representational puzzle: young children can strategically deny and conceal transgressions before they reliably succeed on standard measures of false-belief understanding. Standard interpretations often force a choice between two unsatisfying extremes: either early deceptive behavior is reduced to a routine punishment-avoidance response, or it is taken to imply a surprisingly rich capacity for belief representation. This paper develops an intermediate alternative. Early deception is argued to be better explained by an access-based epistemic policy that tracks another agent's epistemic standing, specifically whether that agent is in a position to know, on the basis of perceptual access, evidential availability, and blocking conditions. This proposal is developed as a policy framework in which actions such as denial, concealment, and withholding are selected under uncertainty as functions of graded epistemic standing. The developmental literature is treated not as the paper's main payload, but as a constraint on representational format. The resulting pattern is asymmetric: young children show flexible sensitivity to witness access, audience knowledge, and opportunities for concealment, yet remain brittle under follow-up questioning, semantic leakage, and evidence-coordinated cover-story demands. This pattern is best understood as evidence for a factive, access-based form of epistemic mindreading that precedes robust belief-based deception.
This work set out to build a strong VGC agent and report what that took, and found that on the live Showdown best-of-three ladder, the agent wins 59% of 150 sets against a human field averaging ${\sim}1320$ Elo.
Evaluating three frontier models across six games in homogeneous and heterogeneous groups over 10 rounds, it is found that different models interpret announcements incompatibly, some as binding commitments and others as cheap talk, producing payoff gaps that emerge in Round~0 and persist across all 10 rounds.
Jerick Shi, Terry Jingchen Zhang, Bernhard Scholkopf et al.· 0 citations
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.
Pranav Aggarwal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.