Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
CoDA is introduced, a fully unsupervised framework that creates reliable privileged information entirely from the latent uncertainty structure of a model's own unlabeled rollouts and provides robust regularization without requiring the strong assumption that the consensus is the absolute ground truth.