One hypothesis is that the kinds of people who feel compelled to immediately become gurus/teachers (rather than just chopping wood, carrying water, chilling, etc.) are more likely to be motivated by a covert egoic desire for power or adulation that leads to hypocritical actions later down the line. Again, PNSE is a catchall term for a wide spectrum of experiential "recontextualizations/belief updates", so the guru may have had some degree of awakening, and even some degree of skill at guiding others to it, while still having skeletons in the back of their ... (read more)
Have you thought at all about whether deliberately inducing "awakening" or a similar kind of ontological crisis might be a useful tool in the toolbox of a "Brainlike AGI Parent" trying to raise a well behaved AGI? (ex. whether it might be useful for establishing a particular predictable regime of reflective stability, in concert with normative social emotions/upbringing, etc.) From my own meditation experience and the anecdotes of other people, it feels harder/more aversive to be "bad-faith" with oneself and others (this includes the ambient "imposture" in... (read more)
An AI without a factorized motivation system (like an LLM) seems more immediately corrigible than an AI with a motivation system, with a lower ceiling for capabilities. Fine-tuned LLMs are capable of accomplishing tasks, and it's possible to anthropomorphize them, but the "value knowledge" upstream of their behaviors exists in an ad-hoc, fractured state, and it can be crudely manipulated. Emergent Misalignment and Subliminal Learning are nice examples. The LLM training process is like an industrialized process for growing a big, connectionist grammar of ab... (read more)
Ha--I was about to type something very similar. The above comment in particular is also ambiguous--it could be alternatively read as coming from someone who has decided that nothing in the LLM reference class will ever be capable enough to be threatening or "incorrigible" in any meaningful sense because of underspecified inductive biases that promote ad-hoc shortcut learning, and maybe intrinsic limitations to human text as a medium to learn representations from. Someone might argue "Aha but by TRAINING the AI to predict the next token or perform longform ... (read more)
Thank you for your response!
Yeah, I guess the challenge for someone arguing in favor would be either (a) operationalizing kindness as an AGI-relevant phenomena and providing evidence based on representative humans who claim to have attained various states of PNSE or (b) coming up with a mechanistic account (compelling enough to persuade a "non-PNSE" person) of how what we colloquially think of as a kind, "virtue-ethical" headspace can emerge naturally and durably from the "right type" of PNSE cultivation (again, in concert with normative human social emoti... (read more)