If we successfully slow down the speed of AI development, but do not perform any 'Science of Alignment' in the meantime, what will we gain? Unless you think an outright global ban on superintelligence is tenable (we will need a global-catastrophe-level warning shot to generate the political will for this IMO, and even then hard for this hold durably), then we have to use the time wisely to advance alignment science.
A reasonable position might be that all prosaic alignment attempts are doomed, so we need to allocate our time / resources to non-prosaic alig... (read more)
Regarding your speculation about the 'WDNYN fns22 E!!q' experiment
But if you actually fine-tune the model on this corpus, I expect it will learn what to do very quickly. Much more quickly than if, say, you'd instead used a distinct (but comparably low-prior-probability) infix for each document.
I agree with your intuition here, but has anyone actually done such an experiment?
I think exploring what kind of persona-related patterns are 'easy' versus 'hard' for the model to learn would be quite interesting to explore, it would indicate some of the inductiv... (read more)
Hello!
I am a physicist making the transition to technical AI safety. I was quite an experienced researcher in physics, focusing on machine learning applications in particle physics (LHC). I am relatively well known researcher in the physics-ML community, I landed a pretty coveted tenure-track scientist job earlier this year which I am now giving up. All to say I have a lot of technical ML / empirical research experience but am new to AI safety.
I had casually followed Effective Altruism for about a decade: regular donor to GiveWell, but not engaged in the ... (read more)
Simply treating the stories as evidence about their author’s dispositions does not explain why the helpful character’s quirk transfers more strongly.
I'm not sure I completely agree with this. Self-identification of the author with one of the characters in a story they wrote is a commonly known phenomena. Since the majority of your experiments are training on chat format, everything in the story can be evidence of the underlying traits of the assistant author.
Very often readers believe or speculate that a character in a story to be a stand-in for the autho... (read more)
I agree. I think the minimal argument for why the current paradigm is perhaps fundamentally flawed is:
Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I'm curious where you would place mech interp in this ontology? Two of its major applications seem to be:
So it seems much more on the 'Science' side, but I am hesitant to label all of it as Science? I guess I could im... (read more)