Skip to content

Author

Jonathan Williams

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Oct 2026

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composit...

Jonathan Williams, E. Tureci, Karthik R. Narasimhan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.