Reward Shaping for Robust Refusal in Small Language Models for Retrieval-Augmented Question Answering
It is shown that instruction-tuned models generate answers even when explicitly prompted to refuse when the answer is not supported by the documents, and Reward Shaping for Refusal and Reasoning (RSRR), a reinforcement learning framework that teaches LMs to reason step-by-step over multiple documents, is introduced.