Skip to content

Author

Patrick Wilhelm

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget. We study the fixed-budget decision problem behind this practice: un...

Patrick Wilhelm, Odej Kao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.