Aug 2026· 1 citation· ⚡ 1 influential· 22 references
MathematicsComputer Science
TL;DR
This work decomposes label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change, and derives task-dependent effective temporal spans.
Abstract
Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $\Theta_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.
The \emph{conditional forecast-revision scale} $\It=\{\Var(\E[X_{t+1}\mid\F_t]\mid\F_{t-1})\}^{1/2}$ measures the history-specific size of the forecast update induced by observing $X_t$. Because it is a conditional second moment built from two unknown conditional means, it is not directly observed. We study which estimator of $\It$ should be used under different structural assumptions and computational budgets. The comparison includes a block bootstrap, a conditional-variance model, a fitted state-space model, two $O(1)$ streaming smoothers, and the forget gate of an already-trained recurrent network. An error decomposition separates one-step-prediction error from conditional-second-moment tracking error. We show that externally tuned lag-only smoothers can be inconsistent when $\It$ changes at the sampling scale, although they attain the usual $T^{-2/3}$ mean-squared-error rate ($T^{-1/3}$ for $\It$) under slow variation; a correctly specified state-space estimator escapes this limit by using the current state. In volatility-driven designs, a cheap conditional-variance model is more accurate and over one hundred times cheaper \emph{as a point estimator} than the implemented block bootstrap, whose value lies in the sampling distribution it provides rather than in point tracking. In state-driven designs, only the structurally matched filter recovers the fast variation. Read directly, a trained network's forget gate does not track $\It$ --- though a supervised linear probe on the full gate vector does, so $\It$ is linearly decodable but not available for free. These results yield a practical rule: identify the conditional-second-moment structure, match the estimator to it, and then choose the least costly adequate method.
A system often has to act long before it learns whether the act worked: a recommender sees a click in seconds and a purchase in days. With $K$ actions and a delay of $d$ rounds, the best rate known for this setting is $\widetilde{O}(\sqrt{(K+d)T})$ over $T$ rounds, so a longer menu is always more expensive to learn from. It need not be: if the outcome depends on the action only through the state it produced, then one late outcome informs every action that could have produced the observed state, and the price is set by how many genuinely different states the actions produce rather than by how many actions there are. We measure this using an effective dimension $v_t$ between $1$ and the number of states, and prove $\widetilde{O}(\sqrt{(d+1)V\log K})$ for a rotating algorithm and $\widetilde{O}(\sqrt{V^{-}}+\sqrt{dT})$ for the single-copy algorithm used in practice, for any budget fixed in advance; merging similar states lowers the price further, at an explicit bias. Even when given the exact losses from $d$ rounds ago, no algorithm escapes $\Omega(\sqrt{dE\min\{1+\log J,T/d\}})$, where $J$ counts the drifting directions and $E$ bounds how far losses move while the learner waits. On generated data, the state channel cuts regret by up to 79 percent against action-level weighting and, on the funnel family, by 32 to 68 percent against a tuned minimax-optimal method.
Latent trait models differ less in what they measure than in what they fix. Item response theory frees the trait from the items administered but, in doing so, surrenders the origin and unit of the scale to convention. The Cognitive Trait Model (CTM; Choi, 2022) restores both by bounding the trait on [0, 1], where 0 denotes an ignorance level and 1 a mastery level defined by the task domain. We generalize the derivation behind CTM into a family indexed by the mastery anchor $L$: the published model is the case $L = 1$, while the limit $L \to \infty$ gives a half-truncated link with an absolute origin but no ceiling. Origin before unit. Bounded support admits Gauss-Legendre quadrature without truncation error, so item and person parameters follow from Bayes modal EM rather than MCMC. On identical data and priors this reproduces the published estimates ($\theta$ correlation 0.9999) while running 35 times faster than a converged random-walk sampler and 73 times faster than NUTS; a single adaptive-testing update costs about five microseconds. Applied to the time axis, the same link becomes a four-parameter growth curve spanning accelerating, decelerating, plateauing and declining trajectories, though not non-monotone ones. The family reduces to a reparameterization of the logistic model only when all items share a slope, so a measurement advantage is possible in principle; three simulations find none, locating the contribution in interpretation and computation rather than in measurement itself. In the LSAT reanalysis, examinees with perfect scores read 67%, 75% or 100% of mastery depending on the estimator -- the criterion-referenced claim is meaningful only with its uncertainty attached, and the coordinate is not in general the proportion of the domain a person can perform.
A (computationally inefficient) adaptive estimator that, so long as $p$ is a mixture of $k$ symmetric log-concave densities, achieves error comparable with the optimal estimator that knows $p$ and has $\tilde\Theta(n/k)$ samples.
This work composition the observables'likelihoods in per-task free-routed last-layer beliefs on a shared backbone absorbs unit-dependent loss scaling into likelihood parameters learned in the same gradient pass, and results land where theory puts them.
This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.
Yi-Fan Zhang, Steve Ta, Jasper Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.