For what it’s worth, I only recently started working full-time in AI safety, but have not encountered stages 6 and 7 frequently when hearing other junior staff articulate the “Constellation worldview.” Instead, I usually hear it articulated as “preventing an AI capable of making significant progress on very ambitious approaches to principled superintelligence alignment from engaging in reward-hacking, sandbagging, etc.,” i.e., stage 4 (but a broader vision of ambitious superalignment than just the agent foundations and CEV stuff), without getting to the st... (read more)
I’ve been wondering: imagine you had a very smart person with very limited ML background and essentially no background in AI safety, but high-potential in some way (e.g., a great philosopher with good mathematical intuitions and very smart). What would be your reading list for them be if they wanted to be an excellent strategic thinker about AI and futurism, and have the context necessary to work on reducing the risk of AI takeover in a productive way?
Some combination of:
That analogy is helpful, thanks. I think I still feel like there’s a gap between developing a good theory of decision-making and putting that theory into immediate practice. For instance, immediately after developing expected utility theory, I could imagine there’s a processing gap among someone who isn’t used to using it to make everyday decisions; I feel a meaningful amount of doubt that von Neumann was much less likely to read books while driving after internalizing expected utility theory (obviously, maybe his preference for doing this was just that st... (read more)
... (read more)I think their plan is more robust than you expect, because "solving agent foundations" is not like "writing down some code", it's more like "totally refactoring your ontology for how rational agents make decisions". From an outside view, having the best understanding of how it's rational to make decisions does seem like it'll very plausibly help you not make obvious mistakes. And from my inside view, I actually think that the solution to agent foundations is specifically action-guiding on questions like "when and how should I try to accumulate power?" (I e
The idea that there are no more good ideas and it's all just scaling from here
Fwiw, this is not how I understand his take. I think he’s saying AI progress will be bottlenecked by compute (and human expert data), which I interpret to mean that the elasticity of substitution between compute/data and ideas isn’t high enough. In fact, I think this his view includes compute to run experiments to do AI research.
I feel like there’s some kind of disanalogy between the prime factorization case (where the goal itself is well-defined (the job is to “find the prime factorization, where N has a process-independent prime factorization”), and hence processes can be refined to become better processes in reaching the goal) and normative goals under the broadly-Humean framework (where the output of the procedure is the goal, and there’s no further goal to reach).
That said, I think there’s something in the vicinity of this that feels right to me. Maybe it has to do with the ... (read more)
It seems to me that:
As a moral anti-realist, I’m sympathetic to something sort of like Humean constructivism, although with a substantial component of “self creation”/understanding that, at the bottom, it’s still on me. In that case, though, I kind of think the values I end up with upon reflection—if the reflection happens in a way I endorse—are what I’d consider my “true values.” This also means that, if the reflection process is underspecified, I get to specify the idealization process I’d like, as it is a process of making, and discovering, myself according to the filters ... (read more)