One pushback re status seeking dynamics
My impression is that, while I can imagine lots of what you write taking courage, my impression is that your contrarian takes have got your lots of status and upvotes and likes
I imagine you have much more respect/status in many ways than you did while at OAI.
You too may be following an incentive landscape
To clarify, I'm thinking bigger picture here.
It's good to have large communities coordinating to do good in ambitious ways
You raise problems that arise when this has been attempted
In the future it would be v good if we could address those problems scalably and sustainably
I get that and it makes sense to focus on the ppl that can do this
But shouldn't you also agree that there's a highly desirable structural piece to make this kind of high integrity approach scalable?
But if ppl went in explicitly saying "we'll differentially advance alignment" but then they mostly accelerate capabilities and start saying "we'll get xrisk ppl in power and that will outweigh the harms"... then I think it's fair to say they get a hit to their credibility and trustworthiness
Like, one accusation Richard could make is "they pessimised their goals overall", where I think it's v unclear for the reasons in the top level comment
But another accusation is "they pessimised their aim of differentially advancing alignment", which I think is pretty p... (read more)
I interpret all Jan's points as agreeing with Richard and saying that "ppl have more integrity" isn't a realistic solution, you need a community/institutions to make it a sustainable equilibrium
I'm surprised Richard doesn't just agree and admit that's a further problem that needs solving
Toby Ord's recent work digs into this. For the case of OAI, i think it suggests most capability improvements are coming from continuing the curve for longer, but there is also some effect from a steeper curve
https://x.com/tobyordoxford/status/1999870642032967987?s=20
However, I'm quite skeptical of this type of consideration making a big difference because the ML industry has already varied the compute input massively, with over 7 OOMs of compute difference between research now (in 2025) vs at the time of AlexNet 12 years ago, (invalidating the view that there is some relatively narrow range of inputs in which neither input is bottlenecking)
Seems like this is a strawman of the bottlenecks view, which would say that the number of near frontier experiments, not compute, is the bottleneck and this quantity did... (read more)
Agree with those updates.
Though a small update as I don't think a default gov-led project would be much better on this front. (Though a well designed one led by responsible ppl could be way way better of course.)
And I've had a few other convos that made me more worried about race dynamics.
Still think two projects is prob better than one overall, but two probbetter than six
I meant at any point, but was imagining the period around full automation yeah. Why do you ask?
I'll post about my views on different numbers of OOMs soon
Sorry, for my comments on this post I've been referring to "software only singularity?" only as "will the parameter r >1 when we f first fully automate AI RnD", not as a threshold for some number of OOMs. That's what Ryan's analysis seemed to be referring to.
I separately think that even if initially r>1 the software explosion might not go on for that long
Obviously the numbers in the LLM case are much less certain given that I'm guessing based on qualitative improvement and looking at some open source models,
Sorry,I don't follow why they're less certain?
based on some first principles reasoning and my understanding of how returns diminished in the semi-conductor case
I'd be interested to hear more about this. The semi conductor case is hard as we don't know how far we are from limits, but if we use Landauer's limit then I'd guess you're right. There's also uncertainty about how much alg progress we will and have met
Why are they more recoverable? Seems like a human who seized power would seek asi advice on how to cement their power
Thanks for this!
Compared to you, I more expect evidence of scheming if it exists.
You argue weak schemers might just play nice. But if so, we can use them to do loads of intellectual labour to make fancy behavioral red teaming and interp to catch out the next gen of AI.
More generally, the plan of bootstrapping to increasingly complex behavioral tests and control schemes seems likely to work. It seems like if one model has spent a lot of thinking time designing a scheme then another model would have to be much smarter to zero shot cause a catastrophe without the scheme detecting it. Eg. analogies with humans suggest this.
I agree the easy vs hard worlds influence the chance of AI taking over.
But are you also claiming it influences the badness of takeover conditional on it happening? (That's the subject of my post)
So you predict that if Claude was in a situation where it knew that it had complete power over you and could make you say that you liked it then it would stop being nice? I think would continue to be nice in any situation of that rough kind which suggests it's actually nice not just narcissistically pretending
But a human could instruct an aligned ASI to help it take over and do a lot of damage
That structural difference you point to seems massive. The reputational downsides of bad behavior will be multiplied 100-fold+ for AI as it reflects on millions of instances and the company's reputation.
And it will be much easier to record and monitor ai thinking and actions to catch bad behaviour.
Why unlikely we can detect selfishness? Why can't we bootstrap from human-level?
One dynamic initially preventing stasis in influence post-AGI is that different ppl have different discount rates, so those with lower discounts will slowly gain influence over time