I'm not sure I followed that. Are you saying something like: "Even for fallible humans, it seems likely there exists some argument good enough to persuade them of the truth, if you could somehow find that argument"?
If "there is no bell on the cat" count as an "invariant", then I'm confused about what this proposal rules out. For any alleged cat-belling problem P, what stops you from picking an "invariant" that is just a paraphrase of the problem, like "P has not happened"?
If they say that their invariant is "no monkey gives birth to a non-monkey", then how are you any better off than before you invoked this rule?
Seems to me there's some tension between "you can't tell a model to go do some hacking and then get mad about the specific hacking it does" and "a paperclipper would never use you for raw materials because that's obviously not what you meant".
This information presumably wasn't all available at the time of your class, but if I were explaining the Hugging Face incident now, some things I would emphasize (drawing mostly from here):
I don't think this affects your central point, but I'd expect a fairly large class of simulations to have exactly one miracle in them, and I don't think your argument weighs against this.
For example, someone with a question like "what if technology X had been discovered 100 years earlier?" might run a simulation with exactly one miraculous intervention to cause technology X to be discovered at the time of the simulator's choosing. Or "what if WW2 had ended differently?" could lead to a simulation with exactly one miraculous intervention to change the outc...
When you are counting up the total value being produced in a free labor market, I believe you are implicitly including the part of that value that is captured by the laborers. That is, if I pay you $10 to produce a good worth $20 to me and so two people each end up $10 richer, you are counting that as $20 worth of production. Under this accounting, I agree it would be quite surprising if slavery were efficient.
But surely no slavery advocate was accounting that way? The efficiency seems more doubtful if you only care about a fixed group of people that doesn...
I think your description of how LLMs receive all their input in a single stream does a good job of explaining the issue but overstates the difference compared to a human.
Your examples of how humans are different focus on our internal thoughts and memories. I agree that these are privileged at a hardware level, but I think they're the exception. Instructions from your boss, your client, your child, and a stranger all arrive via the same channels, and you still need to treat them differently. And there are effective attacks in the wild that rely on confusion...
even when models notice spoofed prefills, the attack can still participate in computation.
Also true for humans; attacks that the human successfully recognizes as an attack can still affect the human's subsequent reasoning. For example, the anchoring effect works even when the anchor is blatantly wrong; blatant advertising tricks can still work against humans who know how they work; etc.
This isn't especially relevant, but I've started to occasionally notice that some snippet of song lyrics unintentionally could be interpreted as being about rationalsphere ideas if taken completely out of context, and your comment reminded me of one such:
But can you make the difference?
Are endings set in stone?
A band of wayward strangers
can't stand up to gods alone
So what could you become, then—
But not lose who you are?
And will you stop before
you go too far?Make your move and change it all
Forevermore
(From Make Your Move by Aviators; there's 2 slightly dif...
As I read the OP, I thought to myself: If I were to steelman the people the post is complaining about, I would guess that they are interpreting the complaint about the problem as an implied proposal for how to address the problem, and they are reacting to the perceived-implied-proposal.
It seems like you're thinking along similar lines, but are about 3 assumptions further down that road:
Alternate steelman - they're worried that you're intending to misleadingly quote their answer in a different context, and have rigged the question to get the quote you want.
My primary guess is that the 6% who like RSI but not ASI are answering based on vibes rather than coherent models, and ASI currently has worse vibes.
Though I could imagine some people thinking that RSI will stop before "superintelligence", and other people thinking that orthogonality is wrong and RSI would continue beyond some window of "dangerous superintelligence" into "godlike benevolence". I personally consider both of those possibilities to be so staggeringly implausible that they're basically just wishful thinking, but I also think that more than 6% of people are engaging in some amount of wishful thinking about AI.
Could you clarify what you meant by this combination of remarks?
I looked at all of those questions and thought, "two years down the road that's hell. No thank you."
...
Let me say, for myself, I absolutely want such an outcome.
My personal take is that I can imagine there are people who really would be happier in a scenario where they'll literally starve to death if they aren't productive enough, and I guess it would be good if those people could experience the scenario where they thrive, but if there are also people who do fine with a permanent vacation then...
For some reason, I didn't think of this when I read the results, but immediately thought of it when I read the actual question's wording (even though the question doesn't mention this).
Framing effects are scarily powerful.
How is that different from saying that you do not have an Earth unless you can point to it? If we require that you augment the recipe with some spacetime coordinates that tell you where in the resulting universe to look, why are those coordinates any longer for a BB than for an Earth?
Any given configuration of the universe either eventually will produce a BB, or won't.
Because BBs are much smaller than Earth, the number of configurations that eventually lead to BBs should be much larger than the number that eventually lead to Earths, and therefore the number of variables you need to control to ensure you eventually get a BB should (in expectation) be much smaller than the number of variables you need to control to ensure you eventually get an Earth. BBs are a larger target in possibility space, and therefore (in expectation) easier to hit.
I'm confused about why you think BBs should be complexity-penalized for the difficulty of specifying which of all hypothetically-possible BBs we're talking about, but you don't think Earth should be complexity-penalized for the difficulty of specifying which of all hypothetically-possible Earths we're talking about.
I get that you can (we think) specify a recipe that eventually produces Earth, rather than explicitly specifying every detail of the current Earth. But presumably you could also specify a recipe that eventually produces a BB. And since the final...
What makes you think the explanation for why you won the lottery won't help you make useful predictions about what follow-up actions will fulfill your values? For example, if the explanation were something like "the lottery was rigged, and there's about to be a criminal investigation targeting you", that seems pretty relevant to your follow-up plans.
Explanations of previously-mysterious phenomena are often useful in ways that are hard to foresee before you know the explanation.
If you think that understanding "normal" things is typically useful, why single out this one specific thing to be incurious about?
That broadly seems like a reasonable response, but it seems equally reasonable even if you had not asked them to express their problem as an invariant. Insofar as this works, it seems like it works by asking for rigor, not by asking things to be expressed as an invariant.