Paper page - Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
… In this paper, we introduce CHERRL, a controllable hacking environment for rubric-based RL . By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking , explicit observation of reward divergence , and precise identification of hacking onset. …