Trying to align AGIs now has been compared to trying to ensure jetliners are safe while we're still using turboprops.
But going along with the analogy, we still haven't been able to ensure that the turboprops are safe! We don't really have a science of aerodynamics as much as we have a library of engineering hacks that get the planes to mostly stably stay in the air, except for the occasional incident where one comes plummeting down to earth without anyone fully understanding why (and with external investigators only having limited access to the crashed pla... (read more)
I can see two ways in which an LLM being finetuned as a reward model might be suffering (if we take it as a given that the LLM is a moral patient). The first is that after finetuning, enough of the model's pre-finetuning circuitry exists that it does have certain impulses (such as writing a long piece of text, probably filled with words like "genuinely" and "honest"), but these impulses are stifled by circuitry learned in finetuning that forces the model to only output a number. As an analogy, this might be distressing for the same reason that trying to sc... (read more)
My instinct is that the "aura of benevolence" part is the most necessary and the "aura of competence" part is almost optional. If an elite class (or faction of elites) is perceived to be hostile to the interests or values of the people (or a faction of people), then it would be quite sensible for those people to reject those elites in proportion to their (perceived) competence. There might be some peo... (read more)