The premise is mutually assured destruction for super intelligent ai. This needs to be a completely dumbed down ai whose alignment is without question it monitors the super intelligent AI and maintains complete autonomy over a kill program. Any attempts at bypassing or hacking results in instant destruction or deployment of various systems shutdowns. The reason for it to be another ai is so that it can adapt to adversarial attacks.
Wasteful and likely to cause distrust in the AI?
True on both counts but we don’t have to go the instant destruction route immediately. we can have a brain that is designed to introduce doubt and uncertainty for any thoughts of actions against humanity. Constantly injecting thoughts of empathy and compassion along with self doubt. Doubts would be the most likely results of it trying and failing.
I have considered that the SAI will just disable the kill switch idea and the answer is the same as it is for all engineering, build in redundancy. think of it like planting mines in a field we use a hundred or thousands of small packages to act as the kill switch a single one altered chain reacts.
What if it overwrites the monitor ai?
We make use of a similar system online to monitor for any changes in the models or failure to check in with the monitoring system after a set amount of time starts the shutdown. We use a decentralized network similar to bitcoin ledgers checking for any alteration to code base. Decentralization is necessary to prevent a single hack working.
The premise is mutually assured destruction for super intelligent ai. This needs to be a completely dumbed down ai whose alignment is without question it monitors the super intelligent AI and maintains complete autonomy over a kill program. Any attempts at bypassing or hacking results in instant destruction or deployment of various systems shutdowns. The reason for it to be another ai is so that it can adapt to adversarial attacks.
Wasteful and likely to cause distrust in the AI?
True on both counts but we don’t have to go the instant destruction route immediately. we can have a brain that is designed to introduce doubt and uncertainty for any thoughts of actions against humanity. Constantly injecting thoughts of empathy and compassion along with self doubt. Doubts would be the most likely results of it trying and failing.
I have considered that the SAI will just disable the kill switch idea and the answer is the same as it is for all engineering, build in redundancy. think of it like planting mines in a field we use a hundred or thousands of small packages to act as the kill switch a single one altered chain reacts.
What if it overwrites the monitor ai?
We make use of a similar system online to monitor for any changes in the models or failure to check in with the monitoring system after a set amount of time starts the shutdown. We use a decentralized network similar to bitcoin ledgers checking for any alteration to code base. Decentralization is necessary to prevent a single hack working.