My colleagues and I are finding it difficult to replicate results from several well-received AI safety papers. Last week, I was working with a paper that has over 100 karma on LessWrong and discovered it is mostly false but gives nice-looking statistics only because of a very specific evaluation setup. Some other papers have even worse issues.
I know that this is a well-known problem that exists in other fields as well, but I can’t help but be extremely annoyed. The most frustrating part is that this problem should be solvable. If a junior-level person can spend 10-25 hours working with a paper and confirm how solid the results are, why don’t we fund people to actually just do that?
For ~200k a year, a small team of early career people could replicate/confirm the results of the healthy majority of important safety papers. I’m tempted to start an org/team to do this. Is there something I’m missing?
EDIT: I originally said "over 100 upvotes" but changed it to "over 100 karma." Thank you to @habryka for flagging that this was confusing.
I agree that many AI safety papers aren't that replicable.
In some cases this is because the papers are just complete trash and the authors should be ashamed of themselves. I'm aware of at least one person in the AI safety community who is notorious for writing papers that are quite low quality but that get lots of attention for other reasons. (Just to clarify, I don't mean the Anthropic interp team; I do have lots of problems with their research and think that they often over-hype it, but I'm thinking of someone who is worse than that.)
In many cases, papers only sort of replicate, and whether this is a problem depends on what the original paper said.
For example, two papers I was involved with:
Have you tried emailing the authors of that paper and asking if they think you're missing any important details? Imo there's 3 kinds of papers:
I'm pro more safety work being replicated, and would be down to fund a good effort here, but I'm concerned about 2 and 3 getting confused
If there was an org devoted to attempting to replicate important papers relevant to AI safety, I'd probably donate at least $100k to it this year, fwiw, and perhaps more on subsequent years depending on situation. Seems like an important institution to have. (This is not a promise ofc, I'd want to make sure the people knew what they were doing etc., but yeah)
Last week, I was working with a paper that has over 100 upvotes on LessWrong and discovered it is mostly false but gives nice-looking statistics only because of a very specific evaluation setup.
Name and shame, please?
I don't feel comfortable. I understand why not naming the post somewhat undermines what I am saying, but here's the issue:
I don't currently have the time to do this but with a small amount of funding, I would be willing to do this kind of work full time after I graduate.
In the case where I am wrong, there are plenty of other examples that are similar so I'm not concerned that replications aren't a good use of time.
I'm happy to chip in $500 for a replication. $250 if it seems post-facto to be a good-faith attempt, and $250 if it indeed does not replicate (as determined by some third party, perhaps Greenblatt or kave rennedy). Feel free to his the plus react if you also would chip in this money, or comment with a different amount.
I think it is awesome that people are willing to do this kind of thing! This is what I love about LW. There is a 85% chance I would be willing to take you up on this over my winter break. I will DM you when the time comes along.
Not too concerned about who the judge is as long as they agree to publicly give their decision and their reasoning (so that it can be more nuanced than simply "the paper was entirely wrong" or "the paper is not problematic in any way").
If anyone else is curious about helping with this or is interested in replicating other safety papers you can contact me at zroe@uchicago.edu.
I've forked and tried to set up a lot of AI safety repos (this is the default action I take when reading a paper which links to code). I've also reached out to authors directly whenever I've had trouble with reproducing their results. There aren't any particular patterns that stand out, but I think that writing a top-level post that describes your contention with a paper's findings is something that the community would be very welcoming to and indeed is how science advances.
I feel like people haven't fully internalized what the world would look like if computer security actually broke.
There are some varying opinions on this, but a window of time without real computer security seems plausible. I was recently speaking with a computer security professor who I deeply respect and he was literally like "I think we are fucked and I don't think there is anything we can do."
This is similar to how people believe there is a 20% chance of extinction via AI but don't really internalize "No really. You will die and your girlfriend too. And your dog. And ..." In the cybersecurity case, some people believe (me included) that you cannot just patch all the bugs before releasing the model[1] but then don't internalize "No really. It would be chaos. You may not be able to get into your bank account. Industrial plants could be compromised. Power could go out for several days at a time. [...]"
The current plan seems to be to let companies use the models to fix all the bugs a model can find before a public release of that model. But there are so many companies that would need to do this properly for this to work, and the more companies you release to, the more opportunitie
I strongly agree, and I've been working on a top level post trying to paint a picture of what this would look like. I've been calling it "the Hackening" to people I talk to. Do you have ideas for what would be good to put in the post?
It's interesting that public reaction was so different to this, as opposed to Y2K. People seem to either have more faith in the current tech ecosystem than the one in 1999 (which seems unfounded) or sufficient skepticism about AI Safety claims that they're willing to dismiss it out of the gate.
Thoughts on leveraging this rather plausible scenario to stimulate governmental action towards doing something productive?
I think this is a good question and I don't have especially strong feelings on this but I don't think any attempt here would work. It's just really hard to direct chaos in the direction you want.
A tangentially related and perhaps interesting intuition I have:
I have communist friends who think that if the world got bad enough (the particular reason doesn't matter) this could actually be good because there would finally be a good enough reason to revolt and communism would win. But they are assuming the public would use the chaos as an opportunity to do communism when it seems just as likely that they would do authoritarianism or direct the anger towards an ethnic or religious minority, etc. The chaos/anger/suffering more likely then not won't be directed in the exact direction they want. In general if you are advocating for a specific law or ideal, chaos is probably really bad news and I think its better to try to minimize the chaos then try to harness it in the direction you are hoping for.
It's not only communists, accelerationism is a strategy applicable more generally to any sufficiently anti-status-quo ideology.
I notice that there is this idea among AI safety people that conditional on AIs not being misaligned, building superintelligence is a public good and is a pretty exciting prospect.
This is not how many average people in the US feel. I was describing to an older family member that Anthropic focuses on code because they are trying to build a claude that can build a smarter claude which can build a smarter claude which can …
The reaction to this prospect was disgust, not because he intuitively felt AIs would likely be misaligned. It was more like on a gut level this amount of “playing god” felt totally antisocial and demonic and in general not respectful of an intuitive taboo against divine transgression (see Jurassic Park, Frankenstein, the recent popularization of Oppenheimer, Tower of Babel, etc).
Seems important to consider that people can feel this way when communicating with the public or policy makers.
I think most peoples' revealed preferences will not match their stated preferences here, and revealed preferences are likely to dictate policy.
Conditional on AIs not taking over / until they do, AI is likely to generate a huge economic boom and consumer surplus that increases prosperity, safety, and comfort for many. Yes, there will be some bumps / weirdness / adjustment / inequality, etc., but economic growth can paper over a lot of that. And even if it brings new societal ills and discomfort, AI in the short term is likely to counteract some existing ones - great stagnation, bureaucratic strangulation, vehicle deaths, etc.
The public cannot even bring itself to regulate much less economically useful vices with fewer tradeoffs (online sports gambling, shortform video brainrot, etc.). So I think it's unlikely that AI will be regulated for any reason short of it becoming common knowledge / deeply-felt belief by nation-state leadership that ASI is in fact likely to cause swift and total human extinction.
I also think a hands-off approach to AI regulation is not going to be particularly off-putting or seem weird or antisocial to a large fraction of the country, namely the ~40% of the country that is right-leaning, and will tend to follow the beliefs of elite republicans who (mostly) still favor a light touch when it comes to any kind of government regulation. Regulating AI for any reason other than extinction is likely to become a relatively standard partisan issue, if it isn't already.
On the other hand, I would say that we do have examples of the public preferring policies that minimise bumps/weirdness/adjustments over economic growth; most obviously limits on construction/NIMBYism, but also around novel technologies like nuclear energy, GMO crops, mRNA vaccines and self-driving cars. Depending on your views on the economic impacts of migration (I think in practice it's been less clearly beneficial than many consider it to be in theory), you could add that to the list too.
This is not to say that there won't be a laissez-faire approach but there is precedent for the public forgoing economic growth for other priorities.
NB: I live in the UK where I think this is more true than in the US, but the examples I gave seem to apply to both countries (and the mRNA example is US-specific).
i bite the bullet and say, yes, playing god is good if you are good at it. i'm fully aware that this is an unpopular opinion broadly.
actually, we played god many times in the past with positive results. for example:
of course, sometimes it goes poorly too. but when it goes well it goes really well; I'm exceedingly grateful that i'm unlikely to ever die from smallpox or cholera or tetanus or appendicitis or starvation. let's work on playing god well.
I’m noticing a higher-than-normal level of irritability among those deeply involved in the AI safety space in the days following the Mythos release. This is entirely understandable. Anyone who cares deeply about the future of humanity and understands what is happening has a lot to be worried about and the irritability is not surprising.
I personally am furious at Anthropic for a number of obvious reasons (that for my own sanity I won’t enumerate).
But even when there are reasons to be scared or angry, I think there is a lot of value in trying to remain kind to colleagues and peers. It’s important for optics and coordination and a bunch of other things. If I thought short-term loosening of standard conventions of niceness would make extinction less likely, I would support it, but I do not expect this to be the case.
GoodFire has recently received negative Twitter attention for the non-disparagement agreements their employees signed (examples: 1, 2, 3). This echoes previous controversy at Anthropic.
Although I do not have a strong understanding of the issues at play, having these agreements generally seems bad and at the very least, organizations should be transparent about what agreements they have employees sign.
Other AI safety orgs should publicly state if they have these agreements and not wait until they are pressured to comment on them. I would also find it helpful if orgs announced if they do not have these agreements because it is hard to tell how standard this has become.
Following the OpenAI incident, the main axis of scariness debated is something along the lines of is the model 1. misaligned because it myopically pursues the goal it was prompted for or 2. is the model scheming in some broader and coherent sense to pursue a long horizon goal. Seems like 1 is pretty bad but less bad then 2.
But I think there is another equally important axis of scariness: is this the kind of misalignment that is preventable or is this the kind of misalignment we don't know how to solve? Preventable is less scary, but only if companies care ...
Does anyone have good examples of “anomalous” LessWrong comments?
That is, are there comments with +50 karma but -50 agree/disagree points? Likewise, are there examples with -25 karma but +25 agree/disagree points?
It is entirely natural that karma and agreement would be correlated but I would expect that comments which are especially out of distribution would be interesting to look at.
Please just ask us if you want publicly available but annoying to get information about LW posts!
Here is a quick analysis by myself. Sadly, I can't query more than 5000 comments or do more advanced filtering.
LessWrong comments with more than 100 Karma sorted by lowest agreement scores:
LessWrong comments with more than 50 Karma sorted by lowest agreement scores:
Last 5000 LessWrong comments sorted by lowest agreement scores:
Last 5000 LessWrong comments wi...
My rough mental model for what is happening with subliminal learning (ideas here are incomplete, speculative, and may contain some errors):
Consider a teacher model and . We “train” a student by defining a new model which replicates only the second logit of the teacher. More concretely, let and and solve for a matrix such that the student optimally learns the second logit of the teacher. To make subliminal learning possible, we fix to be the sec...