>If personas are a viable path to near-term alignment (and I think they are), control could set up a more adversarial relationship with the AI and increase the probability of misalignment that way.
I have some opinions on this (vibe based too):
Partially agree.
If AI development is more insight-driven than compute-driven, then there is more room for sudden progress that gains a decisive advantage over other labs and govs before getting noticed (other entities suspecting the lab getting close to ASI with non-neglible confidence) and reacting. This allows the lab to control the singleton instead of the mainstream labs and govs, and in this situation, the lab might escape from race dynamics.
However, this scenario results in a random lab controlling a singleton. While it's not as hopeless as a singlet... (read more)
https://acoup.blog/2020/03/20/collections-why-dont-we-use-chemical-weapons-anymore/
This article explains it for me: the main reason is effectiveness. Chemical weapons don't work well against specialized protections.
https://www.lesswrong.com/posts/TyusAoBMjYzGN3eZS/why-i-m-not-a-bayesian
In most cases, when we mention probability, we're using bayesian framework. The lack of description power of a scalar probability may be caused by the limits of bayesianism.
While price gouging can quickly mobilize forces to satisfy the emergency demand, they can also have problematic second-order effects. If a price gouge is too high, then this allows certain agents to benefit from the disaster. This creates a perverse incentive that disincentivizes disaster prevention, and potentially even incentivizing artifically intensifying / creating disasters.
For the banning of these weapons, how much does effectiveness weigh against moral concerns? If usefulness weighs a lot, then these examples won't generalize to TAI.
Unless there are very clear, convincing evidence that TAI isn't controllable with current paradigm, then it will still be perceived as a highly useful tech. (Even if such evidence exists, IMO there's high possibility that they'll just cope harder.)
Biochemical weapons: These are only useful against civilians and pre-modern armies. Modern armies can easily afford equipments to protect against the... (read more)
Obedience in this post looks like corrigibility which has been discussed a lot.
I think anti-aging is very different from ai safety in one perspective: it's affected by social wealth distribution. Better ai safety results in AI causing less trouble like Huggingface incident, and less x-risk, so everyone benefits from it. However the anti-aging cure could be expensive, and limited to only the riches and those in charge of political power.
I think this difference will significantly affect how the public views it.