Skip to content

Author

S. Venkatraman

We have 2 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

This work proposes two complementary strategies to improve the performance of value function RL: Privileged Value Functions (PVF) which provide an elegant mechanism to inject additional task-relevant token-level signal without biasing the policy objective; and TETHER, a baseline that adaptively interpolates between group-relative and value baselines depending on the value function accuracy.

S. Venkatraman, Matthieu Dinot, Laurence Aitchison · 0 citations

Amortizing intractable inference in diffusion models for vision, language, and control

Amortized sampling of the posterior over data is studied, and the asymptotic correctness of a data-free learning objective, relative trajectory balance, is proved for training a diffusion model that samples from this posterior, a problem that existing methods solve only approximately or in restricted cases.

S. Venkatraman, Moksh Jain, Luca Scimeca et al. · 75 citations · ⚡5

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.