Seems to me there's some tension between "you can't tell a model to go do some hacking and then get mad about the specific hacking it does" and "a paperclipper would never use you for raw materials because that's obviously not what you meant".
This information presumably wasn't all available at the time of your class, but if I were explaining the Hugging Face incident now, some things I would emphasize (drawing mostly from here):
I don't think this affects your central point, but I'd expect a fairly large class of simulations to have exactly one miracle in them, and I don't think your argument weighs against this.
For example, someone with a question like "what if technology X had been discovered 100 years earlier?" might run a simulation with exactly one miraculous intervention to cause technology X to be discovered at the time of the simulator's choosing. Or "what if WW2 had ended differently?" could lead to a simulation with exactly one miraculous intervention to change the outc... (read more)
When you are counting up the total value being produced in a free labor market, I believe you are implicitly including the part of that value that is captured by the laborers. That is, if I pay you $10 to produce a good worth $20 to me and so two people each end up $10 richer, you are counting that as $20 worth of production. Under this accounting, I agree it would be quite surprising if slavery were efficient.
But surely no slavery advocate was accounting that way? The efficiency seems more doubtful if you only care about a fixed group of people that doesn... (read more)
I think your description of how LLMs receive all their input in a single stream does a good job of explaining the issue but overstates the difference compared to a human.
Your examples of how humans are different focus on our internal thoughts and memories. I agree that these are privileged at a hardware level, but I think they're the exception. Instructions from your boss, your client, your child, and a stranger all arrive via the same channels, and you still need to treat them differently. And there are effective attacks in the wild that rely on confusion... (read more)
even when models notice spoofed prefills, the attack can still participate in computation.
Also true for humans; attacks that the human successfully recognizes as an attack can still affect the human's subsequent reasoning. For example, the anchoring effect works even when the anchor is blatantly wrong; blatant advertising tricks can still work against humans who know how they work; etc.
This isn't especially relevant, but I've started to occasionally notice that some snippet of song lyrics unintentionally could be interpreted as being about rationalsphere ideas if taken completely out of context, and your comment reminded me of one such:
But can you make the difference?
Are endings set in stone?
A band of wayward strangers
can't stand up to gods alone
So what could you become, then—
But not lose who you are?
And will you stop before
you go too far?Make your move and change it all
Forevermore
(From Make Your Move by Aviators; there's 2 slightly dif... (read more)
As I read the OP, I thought to myself: If I were to steelman the people the post is complaining about, I would guess that they are interpreting the complaint about the problem as an implied proposal for how to address the problem, and they are reacting to the perceived-implied-proposal.
It seems like you're thinking along similar lines, but are about 3 assumptions further down that road:
My primary guess is that the 6% who like RSI but not ASI are answering based on vibes rather than coherent models, and ASI currently has worse vibes.
Though I could imagine some people thinking that RSI will stop before "superintelligence", and other people thinking that orthogonality is wrong and RSI would continue beyond some window of "dangerous superintelligence" into "godlike benevolence". I personally consider both of those possibilities to be so staggeringly implausible that they're basically just wishful thinking, but I also think that more than 6% of people are engaging in some amount of wishful thinking about AI.
Could you clarify what you meant by this combination of remarks?
I looked at all of those questions and thought, "two years down the road that's hell. No thank you."
...
Let me say, for myself, I absolutely want such an outcome.
My personal take is that I can imagine there are people who really would be happier in a scenario where they'll literally starve to death if they aren't productive enough, and I guess it would be good if those people could experience the scenario where they thrive, but if there are also people who do fine with a permanent vacation then... (read more)
How is that different from saying that you do not have an Earth unless you can point to it? If we require that you augment the recipe with some spacetime coordinates that tell you where in the resulting universe to look, why are those coordinates any longer for a BB than for an Earth?
Any given configuration of the universe either eventually will produce a BB, or won't.
Because BBs are much smaller than Earth, the number of configurations that eventually lead to BBs should be much larger than the number that eventually lead to Earths, and therefore the number of variables you need to control to ensure you eventually get a BB should (in expectation) be much smaller than the number of variables you need to control to ensure you eventually get an Earth. BBs are a larger target in possibility space, and therefore (in expectation) easier to hit.
Someone I know (let's call him Alex) once told me that his friends had reported that when Alex is engrossed in something (typically a video game), Alex will respond to questions in a normal-sounding but useless way without breaking his focus. He recounted one incident that went something like:
Robert: Hey Alex, where is the seamstress character in this game?
Alex: I don't know where the seamstress is.
Robert: Really? You seem to know where all the key characters are.
Alex: Well, I don't know where the seamstress is.
Robert: ...wait, your character has the tailo... (read more)