PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
It is shown that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning, and introduces PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcem...