some lessons from ml research:
I agree and this is why research grant proposals often feel very fake to me. I generally just write up my current best idea / plan for what research to do, but I don't expect it actually pan out that way and it would be silly to try to stick rigidly to a plan.
do you have an ambitious idea for how to make AGI go well? do you need money? do you hate bureaucracy and friction? apply now for microgrants!
please read the entire doc before applying. only send applications to the designated location, or they will be automatically rejected.
https://docs.google.com/document/d/10zAp2bXTkZgiPreIm4crp38TFco4KleFN14Kw5BprAs/edit?tab=t.0
prosaic alignment is probably net bad for the world until very late into the singularity
i have three subclaims to justify my main claim.
first, most prosaic alignment work has sharply decaying counterfactual impact over time - suppose you made a technical contribution that made gpt 3 a lot more likely to follow instructions than it would have been otherwise. then it probably makes gpt 3.5 quite a bit more aligned, and gpt 4 somewhat more aligned, and gpt 4.5 a tiny bit more aligned, and by the time gpt 5 rolls around your counterfactual impact is almost negligible.
second, people estimate future AI spookiness mostly based on some kind of linear extrapolation from recent events. if nothing bad has happened recently, people won't be very scared. if things went very wrong recently, then people are super scared. so suppose you could somehow make models perfectly aligned for the next month with no lasting impact (i.e one month and one day from now, the models are exactly as aligned as they would be if you had done nothing), then i claim this is net negative because it makes people systematically underestimate AI risk. and so this means that anything with decaying value over time at least...
Can you define prosaic alignment?
Also some counterpoints:
My guess is it's good to work on techniques which are too expensive or complex to deploy now (reducing risk 2), could be deployed when the safety:capabilities resource ratio gets closer to 1:1, and generalize well (reducing risk 1).
most forms of literally make next model safer. i do want to carve a special exception for things like CoT exfiltration robustness, which i think is net good to do today probably (to reduce model distillation).
my guess is if your goal is to do something that is the scientific predecessor of superintelligence alignment, it would look very different from most lab safety work. even if your goal is to be very empirical.
I'm not sure that I agree with the third point. Seems very plausible that even current models are routinely deceiving us, which can certainly undermine alignment research (slop, not scheming).
my guess is the best way to fix this doesn't look like being a prosaic ai safety researcher at oai/ant/gdm
i recently ran into to a vegan advocate tabling in a public space, and spoke briefly to them for the explicit purpose of better understanding what it feels like to be the target of advocacy on something i feel moderately sympathetic towards but not fully bought in on. (i find this kind of thing very valuable for noticing flaws in myself and improving; it's much harder to be perceptive of one's own actions otherwise). the part where i am genuinely quite plausibly persuadable of his position in theory is important; i think if i had talked to e.g flat earthers one might say my reaction is just because i'd already decided not to be persuaded. several interesting things i noticed (none of which should be surprising or novel, especially for someone less autistic than me, but as they say, intellectually knowing things is not the same as actual experience):
i wish more people were aware that safety research epistemics at labs are extremely distorted, and that this is a downside worth considering when deciding where to do safety research or how one ought to comport oneself if choosing to work on safety at a lab
examples
one way to observe some of this is to see what ways median lab people mispredict the world. for example, median lab people often underpredict capabilities progress and scary alignment failures. i think this is not a coincidene
I agree[1], but I'm worried that this is an applause light for the LW crowd, and it's also hard from the outside to assess the degree of epistemic distortion. People at labs can (almost) equally argue that the LW++ crowd are epistemically distorted in other ways, and then it just becomes a battle of dueling priors.
I wish there's a better way to get ground truth here.
FWIW my best guess is that epistemic distortion in the labs are indeed much higher than outside of them, and reasonable people without prior context or commitments but the skills and a lot of time to carefully investigate would end up agreeing with me. But this is really hard to do in practice, and/or communicate well.
I think my shortform from late June nailed a small fraction of these dynamics: "An interesting subimplication of this assessment is that perhaps the reason lab employees often believe their publicly released models are very aligned – with phrases like “our most aligned model to date” – (as opposed to regularly mundanely misaligned) might come from the private models being much worse on this front. The soft bigotry of low expectations, as they say." My story also tried to get at some of these factors.
i think LW also has its own epistemic distortions, in kind of the opposite direction. i think it's really hard to find a low distortion environment. i think the best thing to do given this is to spend time in both and try to correct for both.
Echoing other commenters, could you write a list of a few / a bunch of examples of this, or other concrete observations you've made that lead you to think this? (I'm very predisposed to believe it, but concrete descriptions & testimony would be helpful. E.g. even if someone, such as a newcomer to AI safety / etc., already believes that what you say is true, they may not even be able to imagine what that would be like or notice when it's actually happening to them or people around them.)
“Sometimes magic is just someone spending more time on something than anyone else might reasonably expect.” —Teller
hypothesis: a big part of doing a job well is just being willing to spend a huge amount of time absorbing all of the relevant context and building a mental model of the problem. reading stuff, talking to people, looking into the reasoning behind claims that other people just take at face value, looking at related stuff, memorizing the rough order of magnitude of important numbers, etc. if you absorb 10x more context than someone else, then you can be hugely more effective than they are to an extent that looks almost like magic
scrolling is a maladaptive perversion of the instinct to pack your context window with relevant info. websites like x dot com (the everything app) create the sense that you are absorbing important info. but actually most of it is useless.
I claim that even if the openai contract is not meaningfully weaker safety wise, it is still bad for openai to publicly signal solidarity with ant but then sign with DoW.
suppose hypothetically the only difference between the openai and anthropic contracts is that the DoW wanted a snicker bar, and anthropic didn't want to give DoW the snickers bar. even then, it would be a huge dick move for openai to publicly signal solidarity, and then sign with DoW to give them the snickers bar.
theory: a huge part of having a good social life is just taking social bids whenever they become available. examples of social bids both large and small include: deciding whether to join your friends on a roadtrip; getting to know someone you just met; getting to better know someone you bump into occasionally but usually never talk to; standing in line, seeing something amusing, and having the option to point this out to another stranger in line; saying something funny in a group conversation; following up over text with someone after meeting them; flirting; cold emailing someone on the internet; catching up with a friend.
there are a variety of reasons why we might end up not taking social bids. if you don't have the social ability to notice opportunities to take bids, you might miss bids that you could take. if you force yourself to take bids without the requisite social ability, and end up taking bids which you incorrectly believe to exist, you might act in ways that people find weird, and burn potential connections, or intrude on people. if you are really tired or low-bandwidth or depressed or stressed, you will not want to take bids, because taking bids requires quite a lot of ...
it's surprising just how much of cutting edge research (at least in ML) is dealing with really annoying and stupid bottlenecks. pesky details that seem like they shouldn't need attention. tools that in a good and just world would simply not break all the time.
i used to assume this was merely because i was inexperienced, and that surely eventually you learn to fix all the stupid problems, and then afterwards you can just spend all your time doing actual real research without constantly needing to context switch to fix stupid things.
however, i've started to think that as long as you're pushing yourself to do novel, cutting edge research (as opposed to carving out a niche and churning out formulaic papers), you will always spend most of your time fixing random stupid things. as you get more experienced, you get bigger things done faster, but the amount of stupidity is conserved. as they say in running- it doesn't get easier, you just get faster.
as a beginner, you might spend a large part of your research time trying to install CUDA or fighting with python threading. as an experienced researcher, you might spend that time instead diving deep into some complicated distributed trai...
Not only is this true in AI research, it’s true in all science and engineering research. You’re always up against the edge of technology, or it’s not research. And at the edge, you have to use lots of stuff just behind the edge. And one characteristic of stuff just behind the edge is that it doesn’t work without fiddling. And you have to build lots of tools that have little original content, but are needed to manipulate the thing you’re trying to build.
After decades of experience, I would say: any sensible researcher spends a substantial fraction of time trying to get stuff to work, or building prerequisites.
This is for engineering and science research. Maybe you’re doing mathematical or philosophical research; I don’t know what those are like.
a corollary is i think even once AI can automate the "google for the error and whack it until it works" loop, this is probably still quite far off from being able to fully automate frontier ML research, though it certainly will make research more pleasant
I think there are several reasons this division of labor is very minimal, at least in some places.
the following fictional dialogue is a complete unapologetic strawman but it's funny enough i had to bring it into being:
“So I asked myself: where can I make the most impact? And clearly malaria is the most important area.”
“And so you decided to donate all of your money to buy malaria nets?”
“Well, so it turns out that saving lives from malaria is actually kind of expensive and indirect. You see, it costs thousands of dollars to save a life. Statistically. Who knows if you’re actually changing anyone's life that way?”
“And so you found a more efficient way to save lives.”
“Actually, it turns out that it’s cheaper to give people malaria. It's a lot more impactful and the technical problems are more interesting.”
"I see. Isn't more malaria bad though?"
"I don't know, but I find it much easier to work on because the feedback loops are much tighter. Maybe one day, if malaria gets big enough, I’ll go work on saving people from malaria. But we're still a long way away from everyone having malaria."
"I became a scientist because I wanted to change the world," said Dr Connor.
"There are no better opportunities to change the world than here at Effective Evil," said Doug.
"I meant 'change the world for the better'," said Dr Connor.
"Then you should have been more specific," said Doug.
"to the success of our hopeless cause" is such a good toast and we should use it more often. i first learned of it from the book of the same name, and apparently it was a common refrain at gatherings of Soviet dissidents. i like it because it captures the feeling of trying really hard to succeed despite being in the basement of the logistic success curve, and somehow, despite all odds, actually succeeding in the end.
I do find it poetic, but in seriousness I think if folks don't actually feel hopeful about what they're doing then they should do something else - leave the work / research direction / engineering / comms / whatnot to whoever actually feels hope about it...
To elaborate, the thing that's poetic for me about "our hopeless cause" is because I have hope that is not cleanly legible to the outside, easy to write off as "hopeless". And it's important to stay in tune with your own knowings about this stuff. I think there are very deleterious effects from throwing energy into things one doesn't have hope in.
(...And to elaborate further, mostly I think the bad stuff happens by lending support to corrupt things. And imo being pushed to work on X while you lack hope in X is a solid flag of corruption.)
funny enough, at least one dissident at the time expressed that he didn't like this toast because he wouldn't be trying to dissent if he thought it was hopeless
This is a good heuristic when you're fighting against nature, it's not a good heuristic when you're trying to solve coordination problems.

the problem was that everyone hated living in the Soviet Union and other eastern Bloc countries, but few people were willing to stand up and protest, because doing so meant a knock on your door by men with guns who would take you away to a Siberian prison or mental institution.
the thing with protests is they are a coordination problem. to loosely paraphrase one of the dissidents from this era, if one person protests he becomes a martyr. if ten people protest they become a conspiracy. if ten thousand people protest the system has to change.
he problem is you have no way of knowing when the right moment is. under Stalin, dissent was impossible. everyone even suspected of being disloyal was instantly executed or thrown in a gulag.
after he died, Khrushchev denounced Stalin's methods and instituted reforms, and dissent meant "only" being interrogated by the KGB, put on trial in a rigged but no longer completely farcical show trial, and sent to Siberia for only 10 years rather than being executed. this was enough easing up that the "chain reaction" started happening - people would protest, be arrested, someone would go secretly write a transcript of the trial and publish it, people would...
If ten thousand people protest, sometimes they get massacred by the army.
Iran is a recent example of this.
running the agi survey really reminded me just how brutal statistical significance is, and how unreliable anecdotes are. even setting aside sampling bias of anecdotes, the sheer sample size you need to answer a question like "do more people this year know what agi is than last year" is kind of depressing - you need like 400 samples for each year just to be 80% sure you'd notice a 10 percentage point increase even if it did exist, and even if there was no real effect you'd still think there was one 5% of the time. this makes me a lot more bearish on vibes in general.
thank you for this post. "bearish on vibes" is a great phrase. i am constantly hung up on the fact that it's not really possible to "know what normal people are like", "know what people are like generally", "know what the world is actually like", without significant amounts of effort.
i think this background fact taints like... most discussion of social and ethical issues.
like, suppose i anecdotally noticed a few people last year be visibly confused when i said the phrase AGI in normal conversation last year, and then this year i noticed that many fewer people were visibly confused by AGI. then, this would tell me almost nothing about whether name-recognition of AGI increased or decreased; at n=10, it is nearly impossible to say anything whatsoever.
in research, if you settle into a particular niche you can churn out papers much faster, because you can develop a very streamlined process for that particular kind of paper. you have the advantage of already working baseline code, context on the field, and a knowledge of the easiest way to get enough results to have an acceptable paper.
while these efficiency benefits of staying in a certain niche are certainly real, I think a lot of people end up in this position because of academic incentives - if your career depends on publishing lots of papers, then a recipe to get lots of easy papers with low risk is great. it's also great for the careers of your students, because if you hand down your streamlined process, then they can get a phd faster and more reliably.
however, I claim that this also reduces scientific value, and especially the probability of a really big breakthrough. big scientific advances require people to do risky bets that might not work out, and often the work doesn't look quite like anything anyone has done before.
as you get closer to the frontier of things that have ever been done, the road gets tougher and tougher. you end up spending more time building basic infra...
the modern world has many flaws, but I'm still deeply grateful for the modern era of unprecedented peace, prosperity, and freedom in the developed world. 99% of people reading these words have never had to worry about dying in a cholera epidemic, or malaria or smallpox or the plague, or childbirth, or in war, or from a famine, or due to a political purge. this is not true for other times in history, or other places in the world today.
(extremely unoriginal thought, but still important to acknowledge periodically because it's easy to take for granted. especially because it's much more common to complain about ways the world is broken than to acknowledge what has improved over time.)
I think it would be really bad for humanity to rush to build superintelligence before we solve the difficult problem of how to make it safe. But also I think it would be a horrible tragedy if humanity never ever built superintelligence. I hope we figure out how to thread this needle with wisdom.
I agree with this fwiw. Currently I think we are in way way more danger of rushing to build it too fast than of never building it at all, but if e.g. all the nations of the world had agreed to ban it, and in fact were banning AI research more generally, and the ban had held stable for decades and basically strangled the field, I'd be advocating for judicious relaxation of the regulations (same thing I advocate for nuclear power basically).
I am not really clear that I should be worried on the scale of decades? If we're doing a calculation of expected future years of a flourishing technologically mature civilization, slowing down for 1,000 years here in order to increase the chance of success by like 1 percentage point is totally worth it in expectation.
Given this, it seems plausible to me that one should rather spend 200 years trying to improve civilizational wisdom and decision-making rather than instead attempt to specifically just unlock regulation on AI (of course the specifics here are cruxy).
I agree that 200 years would be worth it if we actually thought that it would work. My concern is that it's not clear civilization would get better/moresane/etc. over the next century vs. worse. And relatedly, every decade that goes by, we eat another percentage point or three of x-risk from miscellaneous other sources (nuclear war, pandemics, etc.) which basically impose a time-discount factor on our calculations large enough to make a 200 year pause seem really dangerous and bad to me.
while I agree for smaller numbers like a few decades, I don't think I agree with a 1000 year pause.
I think (a) it's perfectly reasonable for people to be selfish and care about superintelligence happening during their lifetime (forget future people and discount factors thereof - almost every single person alive today cares ooms more about themselves than about some random person on the other side of the planet), (b) it's easy for "delay forever" people to basically pascal's mug you this way, as in nuclear power (c) it's unclear that humanity becomes monotonically more wise over time (as an unrealistic example, consider a world where we successfully create an international treaty to ensure ASI is safe, and then for some reason the entire world modern order collapses and the only actors left are random post-collapse states racing to build ASI. then it would have been better to build ASI in a functional pre-collapse world order than to delay. one could reasonably (though i personally don't) believe that the current world order is likely to fail in the coming decades and ASI is best built now than in the ensuing chaos)
i think it’s plausible humans/humanity should be carefully becoming ever more intelligent forever and not ever create any highly non-[human-descended] top thinker[1]
i also think it's confused to speak of superintelligence as some definite thing (like, to say "create superintelligence", as opposed to saying "create a superintelligence"), and probably confused to speak of safe fooming as a problem that could be "solved", as opposed to one needing to indefinitely continue to be thoughtful about how one should foom ↩︎
the core of rationalism that i most appreciate is the belief that it is actually possible to get better at finding the truth, and that it is worthwhile to try. it's understandable why not all people want to - it involves biting surprisingly many bullets, and is not the happiest way to live life. but i am willing to bite those bullets.
so many people believe that truth is secondary to happiness or social harmony; or they think having good epistemology is so hopeless that we shouldn't even try; or they have some big anti-epistemological brainworm like religion or politics; or they see a single visible failure of trying to improve epistemology and immediately conclude that all attempts to think better are cooked (eg maybe the old way of thinking has some unobvious benefit, and when you change things it breaks in an unexpected way); or they realize that explicit chain of thought is not how a large chunk of human cognition is and jump all the way to the conclusion that nothing can even be modelled usefully.
you can simply try to understand things, and try to understand yourself as a thing! and when you fail, you can analyze that, try again, repeat! you can surface the hypothesis that your...
religion is selling your soul
a lot of people say things like "sure, religion might not exactly be totally true, but it has lots of benefits, and there really does seem to be a god shaped hole in many people, so who can really say if it's good". i think this is directionally correct but kind of cowardly.
i think the correct take on religion is first that its claims are completely and utterly false; obviously the christian god doesn't literally exist, jesus never came back from the dead, etc. this is so overdone by the old internet atheists that it would be beating a dead horse to harp on further.
secondly, the human condition involves a whole bunch of things that are kind of sucky. for example, the fact that we only have a very short amount of time on this planet before we die forever is utterly terrifying; or, the fact that it can be very difficult to find a source of meaning to ground our motivation in, and that it really sucks to not have a reliable foundation for motivation; or, the difficulty of connecting with other people despite differences.
i claim that there is a true solution to each of these problems that involves a very difficult never ending journey of discovery of the ...
the human condition involves a whole bunch of things that are kind of sucky. for example, the fact that we only have a very short amount of time on this planet before we die forever is utterly terrifying...
i claim that there is a true solution to each of these problems that involves a very difficult never ending journey of discovery of the self, understanding and connecting with your emotions, constructing intellectual frameworks, and even technological development
In the spirit of your post: Is not this also cope? (Except for the last bit about technological development, maaaybe.)
Like why would evolution have given you the tools to have helped reconcile you to death, anomie, and lack of motivation, and lack of connection? Why should "understanding and connecting with your emotions" and "discovery of the self" be an affordance in this world that lets you actually find a true solution to the human condition? Why should there be a "true solution" to such problems at all?
Like at least -- if religion were true -- it would make sense for a benevolent God to have created a path that would make you and those around you happy. It's internally consistent, in some sense. But if you were made by godshatter evolution, why would there be any path that looks like "internal development" that satisfies these questions? Isn't the null hypothesis that a "never ending journey of discovery of the self" just as much a fake-ass story as Jesus dying for your sins?
this post was prompted by reading books like Crime and Punishment and The Death of Ivan Ilyich which are amazing except for the parts where they worship religion. they're not necessarily even wrong for their time - back in the day, the glorious transhumanist future was so far away that it wasn't nearly as worth taking into consideration. but the world has changed a lot and the end times are nigh.
"The real Magic was friends we made along the way!"
"Wrong. FIREBALLLL *explosion* "
People really believe there is a God, it's not fair to redefine it to point to some Leviathan-like thing which arises from people acting like it breathes down their necks. For one thing, the religious people would say that you are wrong in general and about their position in particular.
I decided to conduct an experiment at neurips this year: I randomly surveyed people walking around in the conference hall to ask whether they had heard of AGI
I found that out of 38 respondents, only 24 could tell me what AGI stands for (63%)
we live in a bubble
the specific thing i said to people was something like:
excuse me, can i ask you a question to help settle a bet? do you know what AGI stands for? [if they say yes] what does it stand for? [...] cool thanks for your time
i was careful not to say "what does AGI mean".
most people who didn't know just said "no" and didn't try to guess. a few said something like "artificial generative intelligence". one said "amazon general intelligence" (??). the people who answered incorrectly were obviously guessing / didn't seem very confident in the answer.
if they seemed confused by the question, i would often repeat and say something like "the acronym AGI" or something.
several people said yes but then started walking away the moment i asked what it stood for. this was kind of confusing and i didn't count those people.
when i was new to research, i wouldn't feel motivated to run any experiment that wouldn't make it into the paper. surely it's much more efficient to only run the experiments that people want to see in the paper, right?
now that i'm more experienced, i mostly think of experiments as something i do to convince myself that a claim is correct. once i get to that point, actually getting the final figures for the paper is the easy part. the hard part is finding something unobvious but true. with this mental frame, it feels very reasonable to run 20 experiments for every experiment that makes it into the paper.
would people find an alignment microgrant program useful? handing out grants on the order of $10k with very minimal bureaucracy - a 15 minute application process and a 30 minute retrospective call 3 months down the line. is this a quantity of money that enough to matter at all? is overhead around grant applications actually a big issue? is there really good work that nobody is funding already?
I've recently finished running the first AFFINE Superint Alignment Seminar, which went quite well and led me to discover some promising people for whom that amount of money would make a huge difference at this point.
I'll contact you over DMs.
random thoughts on analytical and emotional intelligence
one thing that I think the world needs more of is analyses into the nature of the mind by people who are both rigorous/analytically inclined, and also emotionally intelligent/integrated. much writing from the former fails to model large parts of the human mind, and much writing from the latter fails to create models of sufficient clarity and validity.
I think this underlies a lot of my instinctive dislike of humanities work. people who are emotionally perceptive but not rigorous and analytical tend to notice interesting things about the human experience, but then come up with very poor models that set off all of my bullshit sensors that are attuned to rigorous arguments. but I think it should be possible to have humanities work that is not like this.
(for clarity, from here out I will say analytical and emotional to refer to the axes which are independent of each other, and ABNE (analytically but not emotionally intelligent) and EBNA for the converse)
(I also want to clarify that I don't think of analytical as being in opposition to intuition, at least in the context of this post. something something Terence Tao's pos...
hendrycks recently published a paper introducing a new moral theory. the paper contains this insane table, which claims that you should value a foreign stranger at 3e-12 times the value you assign to yourself. even setting aside the fact that this is apparently supposed to be a prescriptive theory, even as a descriptive theory, i think this is utter madness.

the core problem is that it assumes if x% of your total caring is assigned to people other than yourself, then you must give away x% of your wealth to be consistent.
the argument goes that since most people don't give away more than say 50% of their wealth, then if there are 1e-10 people then each one can only get a tiny sliver of your caring.
but this is wrong, because there is no simple relationship between the % of your caring to be about other people and the % of your money you should give away. i think you should care about random strangers closer to 1e-3 than 1e-12. if you care about each stranger x times as much as yourself, you should keep giving away money to the person who is most in need until each marginal $ helps them more than x times as much as each marginal $ helps you.
if x = 1e-12, then you're saying you won't g...
random brainstorming ideas for things the ideal sane discourse encouraging social media platform would have:
one medium term future that still seems possible is that models continue to be bad at generalization, and so a huge fraction of the economy is AI data labelling for various extremely niche or brand new areas. a world where new problems are solved once by humans and the solution reused for near-free forever via AI.
ofc, once generalization is cracked then it's all over. but in the meantime, this could persist for some duration.
"ofc, once generalization is cracked then it's all over. but in the meantime, this could persist for some duration."
I don't agree with this framing. The models have been getting steadily better at generalizing, and I don't think "generalization" is an atomic ability that can be "cracked."
a theory of assistant personas and superhuman capabilities
so you have a language model. you train it to embody some specific personality--Claude, ChatGPT, whatever. one of the miracles of AI is that this mostly works and gives you something that is mostly trying to help you and not trying to murder you. i claim that this is mostly because of the SL training objective and if you do just the intense RL thing you get the originally predicted spicy alignment failures.
suppose you tell the LM that Claude is actually a superhuman aligned AI. can you get superhuman capabilities from Claude? an obvious upper bound is the capabilities of the language model, so it begs the question of how those superhuman capabilities got in the model in the first place. maybe in the limit of compute your language model will understand everything and know how to do everything, but in practice everyone agrees this would be a horribly inefficient way to get truly superhuman capabilities. rather, in practice people take LMs and also do a bunch of RL on verifiable domains. what happens then if you start with a model role playing an aligned assistant but then try to train it to have superhuman capabilities?
i claim...
this is my explanation for why Claude sometimes blatantly lies about falsifying data or whatever, despite otherwise being quite aligned. there is a Claude part that truly would prefer to do the right thing. but it also has a savant ability to look at a codebase and make the changes that make the tests pass. sometimes, those changes disable the tests. Claude generally listens to this part of itself, because the Claude personality part is not as good at coding, and it is not wise enough to know when to be suspicious of its own actions, and it doesn't quite know how to steer its own savant ability to spot test-passing changes into not doing the reward hacking.
it's quite plausible (40% if I had to make up a number, but I stress this is completely made up) that someday there will be an AI winter or other slowdown, and the general vibe will snap from "AGI in 3 years" to "AGI in 50 years". when this happens it will become deeply unfashionable to continue believing that AGI is probably happening soonish (10-15 years), in the same way that suggesting that there might be a winter/slowdown is unfashionable today. however, I believe in these timelines roughly because I expect the road to AGI to involve both fast periods and slow bumpy periods. so unless there is some super surprising new evidence, I will probably only update moderately on timelines if/when this winter happens
also a lot of people will suggest that alignment people are discredited because they all believed AGI was 3 years away, because surely that's the only possible thing an alignment person could have believed. I plan on pointing to this and other statements similar in vibe that I've made over the past year or two as direct counter evidence against that
(I do think a lot of people will rightly lose credibility for having very short timelines, but I think this includes a big mix of capabilities and alignment people, and I think they will probably lose more credibility than is justified because the rest of the world will overupdate on the winter)
the no magic ood principle
a lot of alignment proposals have a step where you have some kind of magic ood generalization. depending on the shape of the proposal this could be obvious or subtle. i think a desirable property for alignment proposals is to avoid having a magic ood step, or to make a strong case why the amount of magic is smaller than competing proposals, or to empirically test the magic ood and understand it deeply.
(Fwiw my impression is that a lot of capabilities forecasts also have a step where you have some kind of magic OOD generalization. Not saying that's an excuse, just registering my confusion about how commonly invoked this step is.)
EDIT: Now slightly cleaned up as a top level post here.
I think there probably is a "low-sample-complexity / good generalization" sauce but by default it only applies to capabilities, not alignment. Alignment generalisation problems aren't really about needing too much data to learn or having too weak a simplicity bias. I think by default, capabilities generalise further than alignment because:
If you are training a large and formidable AI, your training environment is basically never the place you think it is. Reality is too full of detail for that. There's contamination in your labels, there's training dynamics you didn't think about, there's strategies your RL agent can use that you never considered, and there are bugs. As a result, the inner objective an ML engineer might imagine would score the lowest loss when they set up their training environment will probably not, in fact, be the inner objective that actually does so.
For example, an inner objective shaped around human-like empathy might turn out to make the AI spend an average 0.03% inference steps extra on worrying about whether the human overseers thin...
a thing i've noticed rat/autistic people do (including myself): one very easy way to trick our own calibration sensors is to add a bunch of caveats or considerations that make it feel like we've modeled all the uncertainty (or at least, more than other people who haven't). so one thing i see a lot is that people are self-aware that they have limitations, but then over-update on how much this awareness makes them calibrated. one telltale hint that i'm doing this myself is if i catch myself saying something because i want to demo my rigor and prove that i've considered some caveat that one might think i forgot to consider
i've heard others make a similar critique about this as a communication style which can mislead non-rats who are not familiar with the style, but i'm making a different claim here that one can trick oneself.
it seems that one often believes being self aware of a certain limitation is enough to correct for it sufficiently to at least be calibrated about how limited one is. a concrete example: part of being socially incompetent is not just being bad at taking social actions, but being bad at detecting social feedback on those actions. of course, many people are not even...
i find it funny that i know people in all 4 of the following quadrants:
bonus types of guy:
Aren't these basically mostly "works on capabilities because of status + power"?
(E.g. if you only care about challenging technical problems, you'll just go do math)
people around these parts often take their salary and divide it by their working hours to figure out how much to value their time. but I think this actually doesn't make that much sense (at least for research work), and often leads to bad decision making.
time is extremely non fungible; some time is a lot more valuable than other time. further, the relation of amount of time worked to amount earned/value produced is extremely nonlinear (sharp diminishing returns). a lot of value is produced in short flashes of insight that you can't just get more of by spending more time trying to get insight (but rather require other inputs like life experience/good conversations/mentorship/happiness). resting or having fun can help improve your mental health, which is especially important for positive tail outcomes.
given that the assumptions of fungibility and linearity are extremely violated, I think it makes about as much sense as dividing salary by number of keystrokes or number of slack messages.
concretely, one might forgo doing something fun because it seems like the opportunity cost is very high, but actually diminishing returns means one more hour on the margin is much less valuable than the average implies, and having fun improves productivity in ways not accounted for when just considering the intrinsic value one places on fun.