Open Internet by MindsNet
Detecting Reward Hacking in Reinforcement Learning
Reward hacking occurs when a policy exploits the reward function in unintended ways, making it difficult to determine if the policy is genuinely improving. This issue arises during the training of reinforcement learning models, particularly when the reward function is complex or imperfect. As a result, it becomes challenging to evaluate the true performance of the policy.
Computing & Technology, Computer Science, Machine Learning