Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the b...
Ming-Yu Chen, Ye-Fan Tao, Gerald Friedland et al.· 0 citations
The Human-LLM Reflection Framework is introduced, a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings, using an information-theoretic analysis based on per-iteration cross-entropy reduction.
The degradation rate across neural models, both sentence embeddings and decoder-only LLMs, is studied, and how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate.