Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers
This work uses teacher-successful problems to define a cheap reference for what the student can learn and measures how the likelihood of each observed token in trajectories from teacher-failed problems changes as an operational learnability signal, which can be computed once from stored trajectories and model checkpoin...