Romain Deléglise
Message
Fundraiser (volunteer) for the french branch of Pause AI. I also have a small blog (in french) about the AI alignment problem.
6
3
Make sense, thanks for the answer
Hello,
Thanks for your post.
I am happy that some people did this research, but I will pushback against some points.
First, inner alignment is not solved even with this results. Because we don't know if the model really internalised the ethical values or some proxy that correlated within the environment. Maybe the new tool (Natural Language Auto encoders) by Anthropic can help distinguish but I don't think the tool is precise and reliable enough.
Second, the generalization seems to be weaker than it appears. Basically the alignment generalized between ethicals...
Maybe I don't understand yours points.
For my point of view we don't need a perfect reward/loss function to be in deep troubles with RSI, that's basically the reason why I expect LLM to likely scale to AGI and then probably ASI.
Likely all we need is a function good enough to begin the circle of RSI. Some domains need humans to verify the quality of the AI answers but it's seems the most important domains for RSI don't need human verification.
This domains are :
- Code (fully automated)
- Maths (fully automated)
- Physic and computer simulations (mostly automated ca
... (read more)