i think some important concepts are being glossed over by using the catch-all "RL training"... but leaving that aside for the more interesting ideas under discussion.
goodhart's law seems relevant here, or, more specifically, 'teaching to the test' and the issues raised regarding the method. I think this article captures the nuance relatively well. Working from the assumption that it is possible to define metrics that accurately assess the alignment of a model, a failure mode of defining metrics for success tends to be optimizing for those metrics, rather t... (read more)
i think some important concepts are being glossed over by using the catch-all "RL training"... but leaving that aside for the more interesting ideas under discussion.
goodhart's law seems relevant here, or, more specifically, 'teaching to the test' and the issues raised regarding the method. I think this article captures the nuance relatively well. Working from the assumption that it is possible to define metrics that accurately assess the alignment of a model, a failure mode of defining metrics for success tends to be optimizing for those metrics, rather t... (read more)