Skip to content

Author

Xiangtian Li

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.

Zhichen Dong, Zhi-Xuan Liu, Yuanjiu Fan, et al. · 0 citations
Preprint Aug 2026

Certified Learning and Equilibrium Implementation under Opaque Partial Commitment

As an extension of existing Bayesian persuasion framework with inadequate message mechanism, we study direct recommendation when a sender is bound by an installed information policy only with probability $\rho$, the realization of binding is hidden, and the receiver does not observe the persistent structural environment. The receiver first sees a payoff-neutral, nonmanipulable calibration sample and then faces a fresh, non-certified deployment interaction. In common, the calibration law identifies only the receiver-facing reduced form, not the latent binding and discretionary kernels. We characterize type-wise $\rho$-implementability, construct the receiver's posterior over the full deployment node, and prove a static direct-following implementation theorem. After every calibration history that passes a posterior-predictive obedience test, the deployment assessment is an exact perfect Bayesian equilibrium: Bayes consistency, receiver sequential rationality, sender sequential rationality, and off-path completion are all verified. Under finite-type separation, common recommendation support, and a positive obedience margin, the test activates such an equilibrium with high probability. Our results keep statistical failure probability distinct from equilibrium approximation. Finally, we embed the original robust value frontier, support-wise linear-programming algorithm, and binary-action fractional-knapsack specialization into this implementation framework

Shuyan Zhang, Xiangtian Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.