Preprint
Jul 2026
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
A controlled experiment on the final window of pretraining, the last data trained on before instruction tuning, finds that what a model is pretrained on last shapes how it reacts to alignment, and what it was trained on last should be reported with it.
Cen Lu, Yung-Chen Tang, Andrea Cavallaro
· 0 citations