I think there's a good argument to be made along the lines that formalization might not only make it so that RSI is secure, but that if it becomes the case that sufficiently advanced AI will optimize itself to be more symbolic-like in order to execute faster (i.e. not have to simulate algorithms in-weight) that formal methods could aid in all three: security (no escape), interpretability (forced-symbolic), and inner alignment (sybolic parts constrained to have certain properties).
Then you'd only need to solve outer alignment.
Do you think the people doing RSI would keep going if there was nobody around like you to help them secure it? Would they let everything get hacked and stolen?
I worry the actual dynamic is: they claim they'll do RSI no matter what, they focus a lot of resources on the dangerous acceleration, because they know that others will step in to do harm reduction.
I worry this is the dynamic of them making a costly threat only because they expect others to respond to it.
wdyt?
I'm also curious if this seems cruxy/important to you (separately from whether you think it's correct)
they claim they'll do RSI no matter what
...
costly threat
? RSI is happening no matter what. just simple commercial competitive pressure even if there weren't national security implications
My motivating example for the morality of working on AI security.
In the early 90s, the decades-long drug corner in Kensington and Allegheny was noticing that people were getting AIDS from sharing needles. In response, the local Act Up chapter got ahold of clean needles and began distributing them. And thus they dubbed the spinoff nonprofit focusing on this Prevention Point, which was promptly targeted by the drug enforcement administration, since needles were illegal for being drug paraphernalia. Many arrests followed by a legal battle later, Philly got a carveout which stands to this day.
Out of the legal battle arose the harm reduction debate. Those in favor of harm reduction say the harm is going to happen anyway so it may as well be less. Those against say that the activists are implicitly condoning the behavior.
I have friends and family who are perplexed that I'm "working on AI" when I claim I do not approve of it. I'm sometimes perplexed as well.
I think they're going to do recursive self improvement (RSI) whether or not I approve. I do not condone RSI, but if its going to happen anyway it might as well be secure.