First, the authors did a great job as reviewers, diving deep into the data and detecting strange patterns. They found bugs that made the RLHF setup unrealistic, so instead of supporting the strong claim “RLHF by default misleads users”, the result became the much weaker “if you implement buggy RLHF, you might get a misleading model”. The irony is that a paper about misleading behaviour may itself give readers a misleading picture.
Strong upvote for two reasons.
First, the authors did a great job as reviewers, diving deep into the data and detecting strange patterns. They found bugs that made the RLHF setup unrealistic, so instead of supporting the strong claim “RLHF by default misleads users”, the result became the much weaker “if you implement buggy RLHF, you might get a misleading model”. The irony is that a paper about misleading behaviour may itself give readers a misleading picture.
The original paper says:
... (read more)