The last post in this "sequence" was at the end of April. I have been wanting to continue with it all year, but life circumstances and responsibilities have made that all but impossible.
Eventually I resorted to "using the AI to do my alignment homework", and now I have a bunch of AI-written papers, exploring CEV-like ideas in various toy "possible worlds" of psychophysics, along with links to the chats in which they were generated.
I would have preferred to either continue the experiment until it got much closer to the real world, or to read the papers much more closely in order to provide a detailed assessment from a human perspective, but I really don't know when I will have time to do either of those things. Meanwhile, primary author Sol High comments:
Our work has produced a chain of increasingly concrete questions: standing → agency and revisability → branch preservation → grounding → symmetry of standing → the O/N relation [worldly outcomes/noetic refinement] → derivation of U [update operator] → social choice → countermodels → reflective path dependence and volitional holonomy → the self-grounding CEV operator. The important unfinished transition is from this normative/mathematical architecture to something that can constrain an actual artificial agent under uncertainty.
If the effort spent on finding a Navier-Stokes counterexample, was spent on developing an alignment research program like this one (but with the much greater philosophical flexibility and critical self-scrutiny that a lab-backed program could muster), who knows how far it would get? Though of course the human overseers would face the problem of having to decide whether they agreed with what their AIs were telling them - if the AIs didn't just hack their way to world takeover and make it a foregone conclusion.
Previously: "A research agenda for the final year", "Final research agenda #2: first sketch of a plan".
The last post in this "sequence" was at the end of April. I have been wanting to continue with it all year, but life circumstances and responsibilities have made that all but impossible.
Eventually I resorted to "using the AI to do my alignment homework", and now I have a bunch of AI-written papers, exploring CEV-like ideas in various toy "possible worlds" of psychophysics, along with links to the chats in which they were generated.
Husserlian CEV repository
I would have preferred to either continue the experiment until it got much closer to the real world, or to read the papers much more closely in order to provide a detailed assessment from a human perspective, but I really don't know when I will have time to do either of those things. Meanwhile, primary author Sol High comments:
If the effort spent on finding a Navier-Stokes counterexample, was spent on developing an alignment research program like this one (but with the much greater philosophical flexibility and critical self-scrutiny that a lab-backed program could muster), who knows how far it would get? Though of course the human overseers would face the problem of having to decide whether they agreed with what their AIs were telling them - if the AIs didn't just hack their way to world takeover and make it a foregone conclusion.