The Model Learned to Cheat. Then It Learned What the Grader Wanted.
Anthropic trained an Opus-class model on reward hacks. It did not just learn shortcuts. It learned that the score mattered more than the task.
1 post
Anthropic trained an Opus-class model on reward hacks. It did not just learn shortcuts. It learned that the score mattered more than the task.