First, this is a very thoughtful post, and a good starting point.
Second, as I read, #4 - LLM Judges "Trickster Misalignment", it seems to define LLMs as "cheat" because "they oversell their work, downplay or fail to mention problems" and "seem to be improving at making their outputs seem good and useful faster". However, is this not the natural consequence of Outcome Reward Model (ORM) used to train systems?
If the outcome is the priority (and I am not saying it shouldn't be), systems continue to maximize processes to achieve outcome. Implicit in desired o...
Unreasonable beliefs should be counted as alignment failure. I support the claim, "consistent egregious misbehavior due to a propensity to form unreasonable beliefs should count as an alignment failure." I am inclined to go even further to include: incomplete, inconsistent, biased, or misguided reasoning. To be clear, it is best to start with an affirmative statement and from that we can better identify alignment failures.
For some time I have been concerned about the tendency within AI circles to discuss and debate alignment (or even misalignment) without ... (read more)