Training a Misaligned Reward Seeker
by evhub, rqi, Monte M, and Benjamin Wright
Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract > During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution...
Sep 123