As your friendly neighborhood transhumanist liberal, I think "trust the science" is mostly a great heuristic. So here's my attempt to trust the science/evidence trans issues. Obviously coming out and transitioning is very scary and very strongly suggests trans people experience something very real, intense and specific and I trust...
TLDR: AI-produced content can reflect random patterns in the data. As such patterns might not get revealed easily, RLHF can accidentally reinforce them. Newer fixes—ties in DPO or in the reward model via BTT, uncertainty-aware RLHF— help us avoid amplifying these quirks but don’t eliminate them. My proposal: explicitly calibrate...
I'm trying to understand how the classical case for AI safety (Bostrom/Yudkowsky argument) best works for LLMs. My impression is that LLMs have passed an important threshold, they have learnt to learn (for instance, GPT 3 understands translation and can learn a new language), however we don't see a rapid...
My model is that 1. Alignment = an AI uses the correct model of interpretation of goals 2. Succeeding in this model design leads to something akin to CEV 3. Errors under this correct model (e.g. mesa-optimisation) are unlikely because high intelligence + a correct model of interpretation = correct...