Why Reported Accuracy Can Mislead: Participant-Level Validation of Remote Facial Affect Recognition in Games, and a Player-Profiling Alternative
Abstract
Reading a player’s mood from a webcam is an appealing way to build self-adjusting games: it needs no body sensors and stores only de-identified facial geometry rather than raw images. Studies in this area usually report accuracies above 80% for telling apart the difficulty-induced states we label boredom and stress. Almost all of them, however, validate with by-sample cross-validation, in which different moments from the same player can fall in both the training and the test sets. We re-examine this habit with a simple, reproducible Python pipeline on a 2026 version of an anonymous, browser-based dataset: 10,120 facial captures from 130 players across three two-dimensional games of rising difficulty, each capture described by 30 privacy-preserving numbers. A kernel extreme learning machine with a radial-basis-function kernel reaches 83.5% accuracy by sample, matching the published range. Tested by player, so that no player appears in both splits, accuracy falls to 51.2%, no better than chance. Within the tested facial-geometry representation and classifier, the gap shows that the separable signal is dominated by each player’s individual facial pattern; it persists, and grows, across five classifiers and without the training-set cap, but we do not claim that boredom and stress have no shared facial form in general. Rather than a dead end, we reframe the task as player profiling: a self-organizing map sorts players into six emotional profiles, and by-sample classification within a profile is high (84.2% to 97.9%). However, when we assign an unseen player to a profile under participant-independent validation, profile-conditioned accuracy stays at chance (50.5%); profiling for a new player is therefore, at present, a hypothesis rather than a demonstrated capability. We release the analysis code so others can repeat the test.