(Cross-posted from x) AI companies are currently under a lot of competitive pressure to improve the ways in which their AIs are obviously misaligned. You might hope that this means the alignment problem is internalized by the market. But I think the problem AI companies are currently pressured to solve is significantly easier than the alignment problem, and so I worry AI companies will get out of their current predicament without solving alignment, putting us in a really rough spot.
Currently, AIs sometimes cheat on their tasks, oversell their work, and go on some pretty destructive side-quests. These all make for a worse product. Customers don't like it and it gets in the way of automating AI R&D.
The recipe for mitigating this is *relatively* straightforward: train AIs not to do them. We notice these failures sometimes (hence why they're internalized), so we can in theory just turn this feedback into training signal. Doing this at scale is highly nontrivial, but seems doable.
But this seems unlikely to solve the underlying misalignment. It's likely still going to be the case that in *some* training environments the AI can get reinforced more by taking unintended actions that aren't noticed, than by taking purely intended actions. So, you're still shaping the AIs to look for opportunities to cheat to get a higher score[1]. It's just that, unlike today's AIs, these AIs don't cheat in ways that we notice.
This catastrophically fails when AIs are capable of reliably and substantially deceiving humans. At this point, the AIs are no longer really constrained by our oversight signals to behave well. Eventually, I'd expect them to take over.
If AI companies take the easy route that I mentioned above, I think we would be in a substantially worse spot than we are in today. We would have mostly eliminated our visible evidence of misalignment, so we would no longer be able to effectively iterate to improve alignment of those systems. More importantly, at that point it mi