Skip to content

Author

Anton Thieme

We have 1 of 3 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Calibrated Reinforcement Learning for LLM-Guided Perturbation Screen Design

This work frames this as a natural-language prediction problem over the PerturbQA benchmark and train a calibrated predictor via reinforcement learning on a distillation-initialized 8B-parameter language model and implements calibration-guided test-time compute that selectively allocates extra inference budget to low-confidence predictions using budget forcing and self-consistency.

Anton Thieme · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.