Progressive Content Refinement with Decaying Reward Joint LinUCB
A novel contextual bandit algorithm that explicitly incorporates reward decay modeling that achieves significant performance gains over strong baselines and confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinement process.