Personally, I believe it would be helpful for the alignment community to somehow quantify how much of a given piece of research goes directly into alignment versus capabilities. But I have recently heard that this task might itself be an alignment-complete problem, which would mean it cannot be solved before the alignment problem itself.
I do not believe that (0.2), but I do not have many arguments. My position is that even though a large part of past progress came from "random" directions unrelated to the final solutions, we do have grantmakers, foundations, safety labs and many other orgs, as well as the personal intentions of the people who decide which research they will actually do, so there should be implicit or (if we do not believe in orgs) at least heuristic arguments available here.
Secondly, there are some existing thoughts on this: post1, post2. By comparison, these arguments personally do not seem persuasive to me: I can imagine cases where some compute research is useless for alignment, and I can imagine the reverse. Everything depends on how hard it is to convert one into the other and what actually will happen with its usage. Can we quantify that, or do we need aligned AGI first?
Personally, I believe it would be helpful for the alignment community to somehow quantify how much of a given piece of research goes directly into alignment versus capabilities. But I have recently heard that this task might itself be an alignment-complete problem, which would mean it cannot be solved before the alignment problem itself.
I do not believe that (0.2), but I do not have many arguments. My position is that even though a large part of past progress came from "random" directions unrelated to the final solutions, we do have grantmakers, foundations, safety labs and many other orgs, as well as the personal intentions of the people who decide which research they will actually do, so there should be implicit or (if we do not believe in orgs) at least heuristic arguments available here.
Secondly, there are some existing thoughts on this: post1, post2. By comparison, these arguments personally do not seem persuasive to me: I can imagine cases where some compute research is useless for alignment, and I can imagine the reverse. Everything depends on how hard it is to convert one into the other and what actually will happen with its usage. Can we quantify that, or do we need aligned AGI first?