Skip to content

Author

Tom Goldstein

We have 2 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

This work introduces $\beta$-OPSD and derives its optimal policy as a geometric interpolation between the reference policy and the privileged teacher, and provides a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.

Jiawei Xu, Minghui Liu, Juzheng Zhang et al. · 1 citation

Not All LLM Reasoning is Visible in the Chain-of-Thought

This work demonstrates a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks and indicates that frontier models already perform consequential computation with no interpretable trace in their output tokens.

Vatsal Baherwani, Tom Goldstein, Ashwinee Panda · 4 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.