The author has some weird misunderstandings about what AI-will-kill-everyone-ism advocates belive, but seems to have a weirdly[1] decent grasp of the problem, given their aforementioned misunderstandings. They argue IRL won't be enough[2]. Here's the interesting quote IMO:
It should be clear that an essential first step toward teaching machines ethical concepts is to enable machines to grasp humanlike concepts in the first place, which I have argued is still AI’s most important open problem.
An example of a weird misunderstanding:
Moreover, I see an even more fundamental problem with the science underlying notions of AI alignment. Most discussions imagine a superintelligent AI as a machine that, while surpassing humans in all cognitive tasks, still lacks humanlike common sense and remains oddly mechanical in nature. And importantly, in keeping with Bostrom’s orthogonality thesis, the machine has achieved superintelligence without having any of its own goals or values, instead waiting for goals to be inserted by humans.Moreover, I see an even more fundamental problem with the science underlying notions of AI alignment. Most discussions imagine a superintelligent AI as a machine that, while surpassing humans in all cognitive tasks, still lacks humanlike common sense and remains oddly mechanical in nature. And importantly, in keeping with Bostrom’s orthogonality thesis, the machine has achieved superintelligence without having any of its own goals or values, instead waiting for goals to be inserted by humans.
For some reason they think "Many in the alignment community think the most promising path forward is a machine learning technique known as inverse reinforcement learning." Perhaps they're making a bucket error and lumping in CIRL with IRL?
mmitchell is a near term safety researcher doing what I view as great work. I think a lot of the miscommunications and odd mislabelings coming from her side of the AI safety/alignment field are because she doesn't see herself as in it, and yet is doing work fundamentally within what I see as the field. So her criticisms of other parts of the field include labeling those as not her field, leading to labeling confusions. but she's still doing good work on short-term impact safety imo.
I think she doesn't quite see the path to AI killing everyone herself yet, if I understand from a distance? not sure about that one.
[This comment is no longer endorsed by its author]Reply
as far as I'm aware, the biggest contribution to the safety field as a whole is mainly improved datasets, which is recent. The sort of stuff that doesn't get prioritized in existential ai safety because it's too short-term and might aid capabilities. In general, I'd recommend reading her papers' abstracts, but wouldn't recommend pushing past an abstract you find uninteresting.
The author has some weird misunderstandings about what AI-will-kill-everyone-ism advocates belive, but seems to have a weirdly[1] decent grasp of the problem, given their aforementioned misunderstandings. They argue IRL won't be enough[2]. Here's the interesting quote IMO:
An example of a weird misunderstanding:
By "weird" I mean "odd for the class of people writing pop sci articles on AI alignment whilst not being within the field".
For some reason they think "Many in the alignment community think the most promising path forward is a machine learning technique known as inverse reinforcement learning." Perhaps they're making a bucket error and lumping in CIRL with IRL?