Related question: wouldn't some findings garner replication-style efforts by default once they become important enough? My sense is that once some finding becomes load-bearing enough (e.g. the METR graph), it inevitably receives critical scrutiny (e.g. critiques of the METR graph). What's the story for why this doesn't happen? Or perhaps it only happens once the paper is past some threshold of notoriety, meaning there's a ton of important but un-replicated papers just below that threshold?
Interesting project! The slow march of AI safety creating an alternate academic universe continues.
I wonder how you're thinking about whether some of your replication is already done within the leading companies. The five criteria you list for judging whether to replicate a paper seem solid, but they also seem to leave out whether the labs might have replicated it already or have a strong incentive to do so.
This is true for a decent amount of empirical safety research (e.g., deliberative alignment, monitoring, etc.), where, if the techniques are effectiv...
I think I’m a little confused about the hypothesis space part. I agree it sounds implausible to run multiple learning algorithms in parallel within a transformer forward pass to find the best one, and the search space is really large.
But if we just ask about the hypothesis space for a moment: is it really practically impossible for a transformer forward pass to simulate a deep-Q style learning algorithm? Even with eg. 3-5 OOMs more compute than GPT-4.5?
I worry you could’ve made this same argument ten years ago for simulating human expert behavior over 8 h...
It feels like the argument of this initiative is: (A) there exist some important safety papers that don't tell the full story, (B) replicating those papers would tell (something closer to) the full story, (C) that type of replication is currently under-incentivized right now, and (D) publishing the full story of those important papers would meaningfully improve safety research. (Tell me if I'm getting it wrong though!)
I buy (A) and (B), but I'm not sure about (C) or (D). I think I'd be more convinced if you had some example(s) from the last 5 years where a... (read more)