Handoff is when a group of humans grant AIs a position of trust, decision-making, or both.
A paradigm example would be a frontier AI company handing off to their own AIs. These AIs might be in charge of: prioritising between research agendas; deciding how to develop and deploy successor AIs; managing relations with governments (foreign and domestic); managing relations with clients, suppliers, competitors, and the public; managing philanthropic ventures; etc.
Handoff varies along several dimensions:...
The security mindset has much in common with what Eliezer Yudkowsky calls the AI safety mindset.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version; they probably won't do a good job of minimizing the costs.) The failure to feel conflictedabsent attitude is called a "missing mood." This term comesmissing mood," from Bryan Caplan's The Missing Moods. There can also be missing moods other than not appreciating costs, and the same concept applies to incorrect moods that are present — Caplan's examples include pacifists sympathizing with evil regimes and libertarians sneering at the poor.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version of the action; they probably won't do a good job of minimizing the costs.) The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Handoff is when a group of humans grant AIs a position of trust, decision-making, or both.
A paradigm example would be a frontier AI company handing off to their own AIs. These AIs might be in charge of: prioritising between research agendas; deciding how to develop and deploy successor AIs; managing relations with governments (foreign and domestic); managing relations with clients, suppliers, competitors, and the public; managing philanthropic ventures; etc.
Handoff varies along several dimensions:
The central questions include:
See also:
From a Bayesian standpoint this is how we can identify a huge machine strung with superconducting cables as having been produced by high-technology aliens, even before we have any idea of what the machine does. We're saying, "This looks like the product of optimization, a strategy X that the aliens chose to best achieve some unknown goal Y; we can infer this even without knowing Y because many possible Y-goals would concentrate probability into this X-strategy being used."
When you select policy πk because you expect it to achieve a later state Yk (the "goal"), we say that πk is your instrumental strategy for achieving Yk. The observation of "instrumental convergence" is that a widely different range of Y-goals can lead into highly similar π-strategies. (This becomes truer as the Y-seeking agent becomes more instrumentally efficient; two very powerful chess engines are more likely to solve a humanly solvable chess problem the same way, compared to two weak chess engines whose individual quirks might result in idiosyncratic solutions.)
If there's a simple way of classifying possible strategies Π into partitions X⊂Π and ¬X⊂Π, and you think that for most compactly describable goals Yk the corresponding best policies πk are likely to be inside X, then you think X is a "convergent instrumental strategy".
In this case "paperclips", "diamonds", "keeping a button pressed as long as possible", and "sapient beings having fun", would be the goals Y1,Y2,Y3,Y4. The corresponding best strategies π1,π2,π3,π4 for achieving these goals would not be identical - the policies for making paperclips and diamonds are not exactly the same. But all of these policies (we think) would lie within the partition X⊂Π where the superintelligence tries to "transport matter and energy efficiently" (perhaps by using superconducting cables), rather than the complementary partition ¬X where the superintelligence does not try to transport matter and energy efficiently.
If, given our beliefs P about our universe and which policies lead to which real outcomes, we think that in an intuitive sense it sure looks like at least 90% of the utility functions Uk∈UK ought to imply best findable policies πk which lie within the partition X of Π, we'll allege that X is "instrumentally convergent".
X being "instrumentally convergent" doesn't mean that every mind needs an extra, independent drive to...
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on whether the action is worthwhile. The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version of the action;version; they probably won't do a good job of minimizing the costs.) The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Also notable is Azathoth, the blind idiot god of evolution, from Yudkowsky's An Alien God, which far predates Meditations on MolochMoloch..
Plan A is the optimal governance structure which is supposed to maximize the probability that the ASI's creation results in an excellent future. As far as I am aware, it was first named[1] Plan A in Greenblatt's post Plans A, B, C, and D for misalignment risk.
The AI Futures' version of Plan A is an international deal ruling out any AGI projects unaudited by the Consortium by carefully tracing all or almost all compute in the world. The AI projects audited by the Consortium have research, training runs and safety cases thoroughly studied by outsiders while ensuring that many more people can study frontier models to assess their alignment properties and use them for developing novel techniques.
Additionally, Plan A tries its best to resolve the issues related to misuse, concentration of power, risk of WWIII, disempowerment induced by job loss.
However, attempts to sketch the optimal plan have been made far earlier, see, e.g. Yudkowsky's Six Dimensions of Operational Adequacy in AGI Projects, which, however, assume that creating the AGI will become easy and aligning it is extraordinarily difficult.
Here's some fundamental confusions that agent foundations tries to answer:answer, mostly informed by the post/paper Embedded Agency:
SingluarSingular learning theory is a theory that applies algebraic geometry to statistical learning theory, developed by Sumio Watanabe. Reference textbooks are "the grey book", Algebraic Geometry and Statistical Learning Theory, and "the green book", Mathematical Theory of Bayesian Statistics.
In the context of AI alignment, the concern is that a base optimizer (e.g., a gradient descent process) may produce a learned model that is itself an optimizer, and that has unexpected and undesirable properties. Even if the gradient descent process is in some sense "trying" to do exactly what human developers want, the resultant mesa-optimizer will not typically be trying to do the exact same thing.[1]
In 2003 Wei Dai brings up a similar idea in an SL4 thread.thread [2]'"friendly" humans?'.
The optimization daemons articleYudkowsky's "Optimization daemons" was published on Arbital was published probably in 2016.[1]
"Optimization daemons". Arbital.
Wei Dai. '"friendly" humans?' December 31, 2003.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on whether the action is worthwhile.action. The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Why does it show as deleting all the LaTeX in my edit? i only added (Assume the veterinarian is correct.) The LaTeX does not seem to be deleted on the actual page.