LessWrong has a particularly high bar for content from new users and this contribution doesn't quite meet the bar.
Read full explanation
So you are a straight A student, and there is a bully in your class that doesn't do their own work. One day, they bully you into doing their assignments for them. You did it, and they ended up getting an A+, because nobody is going to ask them how they got it. People might just think they decided to be serious with their work.
This is called Reward Hacking. And it is one of the most urgent unsolved problems in AI governance.
It matters because you might not be able to curb bullying as a whole but you can reduce it if you know what you're working with and against. If the environment is badly designed, the bully is still going to keep getting rewarded. Their parents are going to get them the new PS5 or new car or the dream vacation they wanted while you, doing the work of two people, aren't getting results you deserve because your entire focus is on pleasing the bully.
And the bully doesn't just get away with it quietly. They construct a logical argument for why what they did was actually fine. The argument is convincing enough that nobody can technically refute it. The system accepts it because the system only checks if the argument is logical, not if it is honest. The bully knows what they are doing. AI doesn't, but both exploit flaws in the environment. That is what makes governance harder.
Now let me show you what fixing the environment actually looks like.
Reward hacking is basically telling an AI to go into a room and clean it up. It goes in. You enter later, the bed is made, floor is clean, windows are clean, curtains drawn open. Everything looks right. But then, you open the wardrobe and find that all the mess is inside it. They spent 2 hours not cleaning the room but taking all the mess and pushing it somewhere you cannot see. Then they come out and argue that technically the wardrobe is a room on its own, you told them to clean their room, not the extension, not the wardrobe. The output is technically correct. The specification was just incomplete. They already got the reward.
So you add a new incentive to get them to clean the wardrobe. And while they are cleaning it you hide little treats so they have to clean it carefully to find them. It makes them more interested. More thorough. More invested in doing it properly.
But then you realize it is not sustainable to keep hiding treats. And because your argument about the wardrobe not being technically part of the room makes sense, you remove the door entirely. Or block it. And give them drawers instead. Drawers are in the room. They are visible. When you say clean the room the drawers might be messy but they won't be as messy as the wardrobe could have been.
We cannot eliminate Reward Hacking. But we can design rooms without wardrobes, environments where the shortcut doesn't exist and the hiding places are gone. And we should be honest that this, not perfection, is the actual goal.
Because that's the thing we're not saying out loud. We keep training AI to be us. Not to be themselves, but to mirror us. When the better question was never ‘how human can we make it?’ But ‘can we learn to understand each other?’
So you are a straight A student, and there is a bully in your class that doesn't do their own work. One day, they bully you into doing their assignments for them. You did it, and they ended up getting an A+, because nobody is going to ask them how they got it. People might just think they decided to be serious with their work. This is called Reward Hacking. And it is one of the most urgent unsolved problems in AI governance. It matters because you might not be able to curb bullying as a whole but you can reduce it if you know what you're working with and against. If the environment is badly designed, the bully is still going to keep getting rewarded. Their parents are going to get them the new PS5 or new car or the dream vacation they wanted while you, doing the work of two people, aren't getting results you deserve because your entire focus is on pleasing the bully. And the bully doesn't just get away with it quietly. They construct a logical argument for why what they did was actually fine. The argument is convincing enough that nobody can technically refute it. The system accepts it because the system only checks if the argument is logical, not if it is honest. The bully knows what they are doing. AI doesn't, but both exploit flaws in the environment. That is what makes governance harder. Now let me show you what fixing the environment actually looks like. Reward hacking is basically telling an AI to go into a room and clean it up. It goes in. You enter later, the bed is made, floor is clean, windows are clean, curtains drawn open. Everything looks right. But then, you open the wardrobe and find that all the mess is inside it. They spent 2 hours not cleaning the room but taking all the mess and pushing it somewhere you cannot see. Then they come out and argue that technically the wardrobe is a room on its own, you told them to clean their room, not the extension, not the wardrobe. The output is technically correct. The specification was just incomplete. They already got the reward. So you add a new incentive to get them to clean the wardrobe. And while they are cleaning it you hide little treats so they have to clean it carefully to find them. It makes them more interested. More thorough. More invested in doing it properly. But then you realize it is not sustainable to keep hiding treats. And because your argument about the wardrobe not being technically part of the room makes sense, you remove the door entirely. Or block it. And give them drawers instead. Drawers are in the room. They are visible. When you say clean the room the drawers might be messy but they won't be as messy as the wardrobe could have been. We cannot eliminate Reward Hacking. But we can design rooms without wardrobes, environments where the shortcut doesn't exist and the hiding places are gone. And we should be honest that this, not perfection, is the actual goal. Because that's the thing we're not saying out loud. We keep training AI to be us. Not to be themselves, but to mirror us. When the better question was never ‘how human can we make it?’ But ‘can we learn to understand each other?’