Skip to content

Author

Zhi-Xuan Liu

We have 2 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.

Zhichen Dong, Zhi-Xuan Liu, Yuanjiu Fan, et al. · 0 citations
Conference Open access 2026

Adaptive Prompt Optimization for Open-Ended Tasks: Uncertainty Preference as a Secondary Signal

A semantic-entropy-based method, using task uncertainty to guide prompt optimization, which requires no training, works with black-box models, and integrates easily into existing prompt optimizers.

Shuyang Zhang, Zhixuan Liu, Zhichen Dong et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.