Calibrated Reinforcement Learning for LLM-Guided Perturbation Screen Design
This work frames this as a natural-language prediction problem over the PerturbQA benchmark and train a calibrated predictor via reinforcement learning on a distillation-initialized 8B-parameter language model and implements calibration-guided test-time compute that selectively allocates extra inference budget to low-confidence predictions using budget forcing and self-consistency.