Archetypal Transfer Learning (ATL) is a proposal by @whitehatStoic for what is argued by the author to be a fine tuning approach that "uses archetypal data" to "embed Synthetic Archetypes". These Synthetic Archetypes are derived from patterns that models assimilate from archetypal data, such as artificial stories. The method yielded a shutdown activation rate of 57.33% in the GPT-2-XL model after fine-tuning. .. (read more)
Religion is a complex group of human activities — involving commitment to higher power, belief in belief, and a range of shared group practices such as worship meetings, rites of passage, etc... (read more)
| User | Post Title | Wikitag | Pow | When | Vote |
The security mindset has much in common with what Eliezer Yudkowsky calls the AI safety mindset.
In the context of AI alignment, the concern is that a base optimizer (e.g., a gradient descent process) may produce a learned model that is itself an optimizer, and that has unexpected and undesirable properties. Even if the gradient descent process is in some sense "trying" to do exactly what human developers want, the resultant mesa-optimizer will not typically be trying to do the exact same thing.[1]
In 2003 Wei Dai brings up a similar idea in an SL4 thread.thread [2]'"friendly" humans?'.
The optimization daemons articleYudkowsky's "Optimization daemons" was published on Arbital was published probably in 2016.[1]
"Optimization daemons". Arbital.
Wei Dai. '"friendly" humans?' December 31, 2003.
Plan A is the optimal governance structure which is supposed to maximize the probability that the ASI's creation results in an excellent future. As far as I am aware, it was first named[1] Plan A in Greenblatt's post Plans A, B, C, and D for misalignment risk.
The AI Futures' version of Plan A is an international deal ruling out any AGI projects unaudited by the Consortium by carefully tracing all or almost all compute in the world. The AI projects audited by the Consortium have research, training runs and safety cases thoroughly studied by outsiders while ensuring that many more people can study frontier models to assess their alignment properties and use them for developing novel techniques.
Additionally, Plan A tries its best to resolve the issues related to misuse, concentration of power, risk of WWIII, disempowerment induced by job loss.
However, attempts to sketch the optimal plan have been made far earlier, see, e.g. Yudkowsky's Six Dimensions of Operational Adequacy in AGI Projects, which, however, assume that creating the AGI will become easy and aligning it is extraordinarily difficult.
SingluarSingular learning theory is a theory that applies algebraic geometry to statistical learning theory, developed by Sumio Watanabe. Reference textbooks are "the grey book", Algebraic Geometry and Statistical Learning Theory, and "the green book", Mathematical Theory of Bayesian Statistics.
Here's some fundamental confusions that agent foundations tries to answer:answer, mostly informed by the post/paper Embedded Agency:
Also notable is Azathoth, the blind idiot god of evolution, from Yudkowsky's An Alien God, which far predates Meditations on MolochMoloch..
Ursula von der Leyen
Mark Carney: p(doom) is "above zero"
Barack Obama: "Nonzero chance" of being wiped out by "killer robots"
António Guterres
Rishi Sunak
Prince Albert II
Naftali Bennett
Ted Lieu (CAIS Signatory)
Audrey Tang (CAIS Signatory)
Bernie Sanders: Bernie Sanders Reveals the AI ‘Doomsday Scenario’ That Worries Top Experts
David Chalmers (CAIS Signatory)
Toby Ord (CAIS Signatory)
Will MacAskill (CAIS(CAIS Signatory)
Others:
Chris Anderson - Dramer-in-Chief, TED (CAIS Signatory)
Lex Fridman (CAIS Signatory)
Grimes (CAIS Signatory)
Stephen Fry (endorsed "If Anyone Builds It, Everyone Dies")
Cenk Uygur: AI's Disturbing Behaviors Will Keep You Up At Night
Paul Tudor Jones: AI poses an imminent threat to humanity in our lifetime
Nate Silver
Steve Bannon
Bill Gates
Kant's thirdsecond formulation of the categorical imperative lets you build up most of the structure of the key moral ideas from a simple rule: "treat no person as purely a means to an end, but always also as an end in themselves". Many applications of the categorical imperative require baroque derivations to loop back and be justified from this premise (treated as a generative axiom) but "consent ethics" in general, and "slavery is forbidden" are both elementary proofs from this starting point. A slave is a person, turned into a tool and piece of property of another person... a literal "means" to ANY end that the owning person (or "Master") deems desirable and feasible.
Historically, slaves would be kept in bondage for life, or for a fixed period of time after which they would be gratedgranted freedom. Many historical cases of enslavement occurred as a result of breaking the law, becoming indebted, suffering a military defeat, or exploitation for cheaper labor; other forms of slavery were instituted along demographic lines such as race or sex.
The issue has gained modern salience given the possible existence of "digital people" who (1) are (1) owned as property, (2) perform cognitive labor, (3) aren't laboring consensually, and (4) don't get pay.
From a Bayesian standpoint this is how we can identify a huge machine strung with superconducting cables as having been produced by high-technology aliens, even before we have any idea of what the machine does. We're saying, "This looks like the product of optimization, a strategy X that the aliens chose to best achieve some unknown goal Y; we can infer this even without knowing Y because many possible Y-goals would concentrate probability into this X-strategy being used."
When you select policy πk because you expect it to achieve a later state Yk (the "goal"), we say that πk is your instrumental strategy for achieving Yk. The observation of "instrumental convergence" is that a widely different range of Y-goals can lead into highly similar π-strategies. (This becomes truer as the Y-seeking agent becomes more instrumentally efficient; two very powerful chess engines are more likely to solve a humanly solvable chess problem the same way, compared to two weak chess engines whose individual quirks might result in idiosyncratic solutions.)
If there's a simple way of classifying possible strategies Π into partitions X⊂Π and ¬X⊂Π, and you think that for most compactly describable goals Yk the corresponding best policies πk are likely to be inside X, then you think X is a "convergent instrumental strategy".
In this case "paperclips", "diamonds", "keeping a button pressed as long as possible", and "sapient beings having fun", would be the goals Y1,Y2,Y3,Y4. The corresponding best strategies π1,π2,π3,π4 for achieving these goals would not be identical - the policies for making paperclips and diamonds are not exactly the same. But all of these policies (we think) would lie within the partition X⊂Π where the superintelligence tries to "transport matter and energy efficiently" (perhaps by using superconducting cables), rather than the complementary partition ¬X where the superintelligence does not try to transport matter and energy efficiently.
If, given our beliefs P about our universe and which policies lead to which real outcomes, we think that in an intuitive sense it sure looks like at least 90% of the utility functions Uk∈UK ought to imply best findable policies πk which lie within the partition X of Π, we'll allege that X is "instrumentally convergent".
X being "instrumentally convergent" doesn't mean that every mind needs an extra, independent drive to...
This looks like the variables are reversed. If V is the intended approximation of U, heavy selection is on high values of V, not U. That selection tends to hit places where V diverges upward from U, so V is an unusually poor approximation of U — not the other way around.