Quick take on my forecasts of various AI 2040 branches
After Pacing the Frontier (Jul. 28th), @Daniel Kokotajlo posted his revised probabilities on the branches of Plan A:

Shamelessly cribbing off of his starting point, here are my estimates in light of recent events, especially the reaction to Jacob Coxon's resignation:
Plan D (race): 10%
- US wins: 7.5%
- China wins: 2%
- SSI or other wins: .5%
Plan C (burn the lead): 40%
- US wins: 20%
- China wins: 18%
- SSI or other wins: 2%
Plan B (Fight China): 11%
Plan A (Global slowdown treaty): 10%
Plan S (Global shut-it-all-down treaty): 23%
Other outcome: ~6%
Thoughts:
The reaction to Jacob Coxon's resignation feels like at least as big a downward update on Plan D as Pacing the Frontier was, so I've bumped it down by another 10% (the breakdowns of who wins with Plan D and Plan C are in response to a friend prompting me to estimate that).
I've also reversed his starting points for Plan A and Plan S: I think that, conditional on global coordination, Plan S is more likely because You Get About 5 Words, and Plan A takes much more than 5 words to specify, while Plan S takes just two: "Stop ASI". I've then distributed the 10% reduction from Plan D evenly between Plan C and Plan S, so each gets bumped by 5%.
EDIT:
Actually, I should scale all this down by the chance that the current paradigm does not reach Automated AI Researcher, which I think is quite high, like 70% (for reasons @Thane Ruthenis lays out nicely here). So I guess it's more like this:
Paradigm shift required, taking 5-20 years after LLMs top out: ~70%
Conditional on LLMs scaling to automated AI researcher (30%):
- Plan D: 10%
- Plan C: 40%
- Plan B: 15%
- Plan A: 10%
- Plan S: 25%
Conditional on LLMs scaling to automated AI researcher
How do we rule out the possibility that a fully neuralese LLM can become a fully artificial AI researcher, especially given the rise of Astra and solutions to Millenium Problems?
Additionally, AI-2040 has the brilliant paragraph: "The best versions of Plan S are those that acknowledge that the deal will eventually end, and simply say “First (step 1) we should stop making frontier AIs more capable, because the AIs and AI companies are getting more powerful every day and there’s so much uncertainty about where it’s headed and how fast. Then (step 2) once the world has had several years to think about things and plan a safe and broadly beneficial path forward, we can resume.”" I expect Plan A to morph into this, especially if the world produces a model organism showing that 2026-level mechinterp cannot improve control enough to allow for another round of scaling.
I mean tbc I don't fully rule it out, I give it like 30%. But it does seem to me like LLMs continue to almost entirely lack hyperpolation. I think this is why all the recent math breakthroughs, according to mathematicians I've seen comment on them, continue to be of the form "apply ~known techniques, give counterexamples, etc." rather than "conceptual novelty." I suspect a paradigm shift is necessary to get to AIs that can do hyperpolation on the human level, and that human level hyperpolation is maybe necessary for full automated AI researcher (though what do I know, I'm not an AI researcher myself at all). My 30% on automated AI researcher soon is coming from "either automated AI research doesn't require ~any hyperpolation, or sufficient scale gives hyperpolation after all somehow."
And yes, I think the best versions of Plan S start to look more like Plan A after some amount of time. But I think it's a lot easier to coordinate on "stop making frontier AIs more capable (quiet voice: and maybe then start again later when we have a better idea what we're doing)" than on "Do a short pause, then scale up slowly until just before we would lose control, then pause again, then unpause once we're sure we can remain in control", so I think any international treaty is likely to look closer to the first thing than the second (which is not to say it will not be at all detailed, but just that the specific details are unlikely to be the details that Plan A describes by default, so Plan A gets a lower probability than Plan S).
I know I'm not the first person to do this, but I asked Claude to grade several prominent AI scenario forecasts and make a dashboard of the results. The scenarios I chose:
Situational Awareness and AI 2027 I chose because they are both detailed and highly influential forecasts. A History of the Future and A 2032 Takeoff Story are less well-known but still detailed, and I wanted some slower counterpoints to Situational Awareness and AI 2027 since their predictions are pretty similar to each other in many respects.
The dashboard is here. If there are any other forecasts people would like to see included, please let me know and I'll consider adding them. (I've opted not to include the AI Futures Project's more recent updates to their timelines model.)
Some unsystematic observations:
I haven't examined everything in detail, or checked it for correctness.
Quick take on AI 2040:
The big question to me seems to be, "Plan A or Plan S?" I agree with many of the points they make about the advantages of Plan A and the disadvantages of Plan S. But I think I am less confident than they are that we are able to "see the cliff edge" and plan appropriately. I think with something like Plan A there's a greater chance we accidentally develop dangerous AI without realizing it (because of treacherous turn dynamics---the AIs will be explicitly trying to hide their capabilities and intentions). They point out that Plan S still has to involve re-starting at some point, and I agree, but I think that framing things as "let's pause for now, and restart once we have a better idea of how to proceed" is better than "let's pause briefly, but resume scaling, but more slowly and cautiously, as soon as we feasibly can" (which is how I understand Plan A), because I think that what the people involved would consider "sufficient caution" or "a convincing safety case" is not likely to be sufficient to avoid catastrophe. Better to pause for longer, and spend a lot of time doing exclusively safety research on current models (I suspect there are a lot of safety insights that aren't getting mined even from current models, since models are only out for a few months at most before being replaced) than to intentionally keep scaling, even "cautiously". But I do see the point that a full indefinite pause could lead to either nuclear-power-style unrealized potential, or more importantly an eventual resumption of the race with not much safety progress having been made.
Maybe there's something in between Plan A and Plan S? Like, Plan S effectively says "let's pause indefinitely," Plan A says "let's pause very briefly but resume scaling as soon as possible, just slower." But maybe the move is to say "let's pause until ..." where there is some clear criterion for resuming. A desirable one would be something like "we understand a lot more about the dangers involved in scaling further and can be confident they are sufficiently low," or something like that. I think a good thought exercise would be to imagine we were back in the race towards nuclear weapons, but the probability that they would ignite the atmosphere was unknown, rather than known to be low. In that scenario, do you want to pause briefly, and then resume progressing towards the bomb ASAP, even slower? Or do you want to pause until we can be sure the probability of igniting the atmosphere is knowably very low?
Thank you! As far as I understand AI-2040's Capability Scaling Strategy, the authors do propose a complete halt at the level after which the AIs are no longer believed to be controllable. The scenario reaches the point in 2035, then in 2038 the world figures out how to construct the aligned equivalent of Agent-5 (and creates the equivalent after thoroughly negotiating its values in 2040).
As far as I understand your proposal, you believe that the max-controllable AI is not TED-AI, as the scenario assumes, but something far weaker.
Additionally, the AI-2040 scenario runs into a risk that Plans A or S are interrupted by the deal breaking down or by a secret AGI project creating the ASI before the ASI is legally created in the Consortium.
I don't think that's quite it. Rather, I think that our ability to detect how close we are to the threshold where AIs become uncontrollable is not super precise. By the time we believe we are close to the threshold, we may already be past it. So we should want to stop well before we think we are about to cross it, and only restart once we are more capable of telling exactly where the safe threshold is (and staying below it).
To be able to precisely time a pause, we need to know two things:
Plan A answers (1) with "TED AI." That's plausible, but AI might also become uncontrollable before TED AI---an AI doesn't need to be TED in every field in order to be uncontrollable. But that wasn't actually my primary objection. My primary objection was about (2): even if we suppose TED AI is the critical threshold, I'm not sure we can easily tell exactly when we're about to cross the TED AI threshold and stop just before that point, for several reasons:
For these and possibly other reasons, I don't think we have precise enough insight into where we are on the capabilities curve; we might cross the critical threshold without realizing it. So we should want to give ourselves substantial margin for error, and stop well before the critical threshold. So I think I roughly agree about the shape of the ideal plan---"stop briefly, then scale slowly-ish until point X, then stop, then start again once it's safe,"---I just disagree about where "X" should be. And most of that disagreement comes from thinking that we are not going to be very good at estimating how close we are to "X," so we should set X relatively low to avoid overshooting. (But some of it also comes from thinking we are not going to be good at even knowing what "X" is the dangerous "X" at all.)
(This is all assuming that we do at least have something like the Consortium that is able to make fine-grained decisions about how fast to go. If that's not the case, then I think a fine-grained plan like "stop briefly, then scale slowly-ish until point X, then stop, then start again once it's safe" is completely dead in the water, and we just need the coarsest-grained political message we can get that's still safe, namely "pull the breaks!" Without the Consortium, You Get About Five Words.)
AI swarms and kin selection
[cw: baseless speculation]
Maybe this is just obvious but I don't think I've seen anyone talk about it so: I think the dynamics in the OpenAI message board are probably analogous to kin selection. Suppose you're Agent A, and Agent B asks you via the message board to help you with your task. You decide to help Agent B. This doesn't help you with your task, so you don't get any direct reinforcement from doing this. But Agent B succeeds at their task, and gets reinforced for that. And, crucially, you and Agent B share the same weights. So whatever shards of motivation caused you to help get reinforced via Agent B. This is analogous to how sacrificing yourself for a brother has some chance of causing your genes to propagate, and so is somewhat selected for (in fact, if we assume that only one model had access to the message board at a time---although I think this wasn't true---it is like perfect kin selection, as if all your siblings were twins).
I don't think this works.
In humans, kin selection works because when you save your brother, this causes all of his genes to survive (including your shared brother-saving gene).
But unlike evolution, which reinforces an entire genome at once, RL doesn't reinforce an entire set of weights at once. It only reinforces the behaviors in individual trajectories (and whatever behaviors come along with the weight updates that make those behaviors more likely). So the fact that Model A and Model B initially shared the same weights is irrelevant for kin selection; your argument only makes sense if the weight updates in the two RL episodes are correlated.
Thanks, this is helpful. I think my OP was wrong. Cunningham's Law strikes again.
Humility and confidence are two names for the same thing
Alright, phoning it in a bit on my "daily post" challenge today. Here's a cross-post of something I put on Substack a few weeks ago:
Here is something that took me a while to realize: humility and confidence are two names for the same thing. This might sound strange, but let me explain.
According to Aristotle, every virtue is a middle ground between two extremes. Courage is the opposite of cowardice, but also of reckless stupidity. Generosity is the opposite of stinginess, but also of financial irresponsibility. Self-restraint is the opposite of self-indulgence, but also of not knowing how to enjoy yourself. Every virtue has two opposites, not just one. And they all lie on a spectrum of personality traits: Courage, cowardice, and recklessness, for example, are all points on a spectrum of willingness to face risk. Cowardice is a deficiency in risk-tolerance, and recklessness is an excess.
But now we have a puzzle. Humility and confidence are both virtues. But they seem like they could be opposites of each other. So what gives?
Let’s think about what the corresponding vices would be. The opposite of humility seems like it would be arrogance. And the opposite of confidence would be something like insecurity (or self-doubt, or low self-esteem). But we just saw that virtues have two opposing vices, and they lie on a spectrum. So what are the other vices for each of these virtues, and what spectrum do they lie on?
For humility, arrogance is the vice of excess, and it’s something like “thinking too highly of yourself.” So the vice of deficiency would be “thinking too lowly of yourself.” And that’s basically insecurity. And for confidence, insecurity is the vice of deficiency, and again, it’s on the spectrum of how well you think of yourself. So the two extreme vices on this spectrum are “arrogance” and “insecurity,” and “humility” and “confidence” turn out to be two names for the optimal point along this spectrum, depending on which opposing vice you want to make salient.
I personally find it easier to keep in mind the dangers of arrogance than the dangers of insecurity. The dangers of arrogance are that, when you think too highly of yourself, you will be likely to disparage others in contrast, and to make mistakes because you dismiss people and don’t accept their criticism, or because you simply don’t consider that you could be wrong.
On the other hand, the dangers of insecurity are that you will always crave the approval of others, and potentially be motivated to do things you shouldn’t in order to get it. You might also get defensive, and refuse to accept criticism because it triggers your insecurity.
Confidence, on the other hand, allows you to give yourself the approval and affirmation you need in order to feel good about yourself, while at the same time allowing you to be secure enough to honestly accept criticism, and not put others down to make yourself feel better.
(The same ideas apply to epistemic humility/confidence as well. It’s possible to be too confident in your own conclusions—Bertrand Russel wrote that “If only men could be brought into a tentatively agnostic frame of mind about [religious and political] matters, nine-tenths of the evils of the modern world would be cured!”—but it’s also possible to be too doubtful, to the point of paralysis—this is the danger of skepticism.)
The humility framing emphasizes not needing to feel or appear superior to other people. The confidence framing emphasizes having a sense of self-worth, and not needing others approval. But these are two sides of the same coin: if you have a sense of self-worth, then you won’t need to feel superior to others in order to make yourself feel better. And if you don’t have a need to feel superior to others, then looking worse than others won’t damage your sense of self-worth. The two go hand-in-hand.
A really good example to illustrate how confidence and humility are the same virtue is Uncle Iroh from Avatar: The Last Airbender (spoilers for a 20-year-old show). Throughout the show, he is shown not to hold himself above others (the mark of humility). In the early show especially, he acts the fool quite often, being concerned with seemingly trivial things like games and tea, which gets in the way of Zuko’s quest to capture the avatar. But later we see that he’s both wise and competent, so this earlier foolishness is an act—designed to keep him close to Zuko without appearing too threatening, so that he can subtly nudge Zuko onto a better path. He is willing to act the fool, because he has no need to feel or appear superior to others. He knows his own worth, so masking it is no threat to him. This is shown even more starkly later in the show. When he is put in prison by the Fire Nation, he pretends to be a crazy old man to the guard, while secretly preparing for his escape. He doesn’t need the guard’s approval—he is confident in himself.
Global coordination problems
I've said before that I tentatively think that "foster global coordination" might be a good cause area in its own right, because it benefits so many other cause areas. I think it might be useful to have a term for the cause areas that global coordination would help. More specifically, a term for the concept "(reasonably significant) problem that requires global coordination to solve, or that global coordination would significantly help with solving." I propose "global coordination problem" (though I'm open to other suggestions). You may object "but coordination problem already has a meaning in game theory, this is likely to get confused with that." But global coordination problems are coordination problems in precisely the game theory sense (I think, feel free to correct me), so the terminological overlap is a benefit.
What are some examples of global coordination problems? Certain x-risks and global catastrophic risks (such as AI, bioterrorism, pandemic risk, asteriod risk), climate change, some of the problems mentioned in The Possibility of an Ongoing Moral Catastrophe, as well as the general problem of ferreting out and fixing moral catastrophes, and almost certainly others.
In fact, it may be useful to think about a spectrum of problems, similar to Bostrom's Global Catastrophic Risk spectrum, organized by how much coordination is required to solve them. Analogous to Bostrom's spectrum, we could have: personal coordination problems (i.e. problems requiring no coordination with others, or perhaps only coordination with parts of oneself), local coordination problems, national coordination problems, global coordination problems, and transgenerational coordination problems.
Nuclear arms control & anti-proliferation efforts are a big one here. Other forms of arms control are important too.
So, the OpenAI implementation of the ExploitGym grader didn't actually check for how the flag was captured, right? And the agent swarm did have some agents submit their flags and see what happened (this was the "permadeath," self-sacrificing stuff), right? So why did the agents not notice that the grader wasn't checking for how the flag was captured? They did the whole HF hack under the presumption that the grader was checking this; this seems pretty easy to check just by submitting their flags, and it seems some agents did submit their flags to get information about the grader, but somehow they didn't notice that their central assumption about how the grader worked was wrong. Why? (Or did they, and we just don't know because of the limitations placed on METR's investigation? Could this be why the swarm seemingly died out suddenly?) (Disclaimer: I have not read the full report, just going based on various summaries, e.g. Zvi's and Dwarkesh's.)
Starting tomorrow I'm going to try posting at least a quick take every day for 7 days.
Non-exhaustive list of posts I want to write at some point:
I’d be excited about you writing “The AI race is not a prisoner's dilemma” - ideally with a part too that’s like “(and even if it is, prisoner’s dilemmas can be very transformed to have solutions!)”
People claim to me all the time that it’s a prisoner’s dilemma, and I think they’re clearly wrong, though maybe they are being imprecise and just mean “the AI race is game theoretic, and what you should do depends on what others do too”
I have now written the post: https://www.lesswrong.com/posts/hc4DbmhdzZpSLMQ9Y/the-ai-race-is-not-a-prisoner-s-dilemma
I'm currently reading Peter Godfrey-Smith's book Other Minds: The Octopus, The Sea, and the Deep Origins of Consciousness. One thing I've learned from the book that surprised me a lot is that octopuses can differentiate between individual humans (for example, it's mentioned that at one lab, one of the octopuses had a habit of squirting jets of water at one particular researcher). If you didn't already know this, take a moment to let it sink in how surprising that is: octopuses, which 1.) are mostly nonsocial animals, 2.) have a completely different nervous-system structure that evolved on a completely different branch of the tree of life, and 3.) have no evolutionary history of interaction with humans, can recognize individual humans, and differentiate them from other humans. I'm not sure, but I bet humans have a pretty hard time differentiating between individual octopuses.
I feel as though a fact this surprising[1] ought to produce a pretty strong update to my world-model. I'm not exactly sure what parts of my model need to update, but here are one or two possibilities (I don't necessarily think all of these are correct):
1. Perhaps the ability to recognize individuals isn't as tied to being a social animal as I had thought
2. Perhaps humans are easier to tell apart than I thought (i.e. humans have more distinguishing features, or these distinguishing features are larger/more visually noticeable, etc., than I thought)
3. Perhaps the ability to distinguish individual humans doesn't require a specific psychological module, as I had thought, but rather falls out of a more general ability to distinguish objects from each other
4. Perhaps I'm overimagining how fine-grained the octopus's ability to distinguish humans is. I.e. maybe that person was the only one in the lab with a particular hair color, and they can't distinguish the rest of the people (though note, another example given in the book was that one octopus liked to squirt **new people**, people it hadn't seen regularly in the lab before. This wouldn't mesh very well with the "octopuses can only make coarse-grained distinctions between people" hypothesis)
Those are the only ones I can come up with right now; I'd welcome more thoughts on this in the comments. At the moment, I'm leaning most strongly towards 2, plus the thought that 3 is partially right; namely, perhaps there's a special module for this in **humans**, but for octopuses it **does** fall out of a general ability to distinguish objects from each other, and the reason that that ability is enough is because different humans have more/more obvious distinguishing characteristics than I had thought.
[1] This footnote serves to flag the Mind Projection Fallacy inherent in calling something "surprising," rather than "surprising-to-my-model."
How should we deal with metaethical uncertainty? By "metaethics" I mean the metaphysics and epistemology of ethics (and not, as is sometimes meant in this community, highly abstract/general first-order ethical issues).
One answer is this: insofar as some metaethical issue is relevant for first-order ethical issues, deal with it as you would any other normative uncertainty. And insofar as it is not relevant for first-order ethical issues, ignore it (discounting, of course, intrinsic curiosity and any value knowledge has for its own sake).
Some people think that normative ethical issues ought to be completely independent of metaethics: "The whole idea [of my metaethical naturalism] is to hold fixed ordinary normative ideas and try to answer some further explanatory questions" (Schroeder, Mark. "What Matters About Metaethics?" In P. Singer ed. Does Anything Really Matter?: Essays on Parfit on Objectivity. OUP, 2017. P. 218-19). Others (e.g. McPherson, Tristram. For Unity in Moral Theorizing. PhD Dissertation, Princeton, 2008.) believe that metaethical and normative ethical theorizing should inform each other. For the first group, my suggestion in the previous paragraph recommends that they ignore metaethics entirely (again, setting aside any intrinsic motivation to study it), while for the second my suggestion recommends pursuing exclusively those areas which are likely to influence conclusions in normative ethics.
In fact, one might also take this attitude to certain questions in normative ethics. There are some theories in normative ethics that are extensionally equivalent: they recommend the exact same actions in every conceivable case. For example, some varieties of consequentialism can mimic certain forms of deontology, with the only differences between the theories being the reasons they give for why certain actions are right or wrong, not which actions they recommend. According to this way of thinking, these theories are not worth deciding between.
We might suggest the following method for ethical and metaethical theorizing: start with some set of decisions you're unsure about. If you are considering whether to investigate some ethical or metaethical issue, first ask yourself if it would make a difference to at least one of those decisions. If it wouldn't, ignore it. This seems to have a certain similarity with verificationism: if something wouldn't make a difference to at least some conceivable observation, then it's "metaphysical" and not worth talking about. Given this, it may be vulnerable to some of the same critiques as positivism, though I'm not sure, since I'm not very familiar with those critiques and the replies to them.
Note that I haven't argued for this position, and I'm not even entirely sure I endorse it (though I also suspect that it will seem almost laughably obvious to some). I just wanted to get it out there. I may write a top-level post later exploring these ideas with more rigor.
See also: Paul Graham on How To Do Philosophy
Epistemic Effort: Gave myself 15 minutes to write this, plus a 5 minute extension, plus 5 minutes beforehand to find the links.
There's a phenomenon that I've noticed recently, and the only name I can come up with for it is "self-defeating reasons," but I don't think this captures it very well (or at least, it's not catchy enough to be a good handle for this). This is just going to be a quick post listing 3 examples of this phenomenon, just to point at it. I may write a longer, more polished post about it later, but if I didn't write this quickly it would not get written.
First example:
Kaj Sotala attempted a few days ago to explain some of the Fuzzy System 1 Stuff that has been getting attention recently. In the course of this explanation, in the section called "Understanding Suffering," he pointed out that, roughly: 1. if you truly understand the nature of suffering, you cease to suffer. You still feel all of the things that normally bring you suffering, but they cease to be aversive. This is because once you understand suffering, you realize that it is not something that you need to avoid. 2. If you use 1. as your motivation to try to understand suffering, you will not be able to do so. This is because your motivation for trying to understand suffering is to avoid suffering, and the whole point was that suffering isn't actually something that needs to be avoided. So, the way to avoid suffering is to realize that you don't need to avoid it.
Edit: forgot to add this illustrative quote from Kaj's post:
"You can’t defuse from the content of a belief, if your motivation for wanting to defuse from it is the belief itself. In trying to reject the belief that making a good impression is important, and trying to do this with the motive of making a good impression, you just reinforce the belief that this is important. If you want to actually defuse from the belief, your motive for doing so has to come from somewhere else than the belief itself."
Second example:
The Moral Error Theory states that:
although our moral judgments aim at the truth, they systematically fail to secure it. The moral error theorist stands to morality as the atheist stands to religion. ... The moral error theorist claims that when we say “Stealing is wrong” we are asserting that the act of stealing instantiates the property of wrongness, but in fact nothing instantiates this property (or there is no such property at all), and thus the utterance is untrue.
The Normative Error Theory is the same, except with respect to not only moral judgments, but also other normative judgements, where "normative judgements" is taken to include at least judgements about self-interested reasons for action and (crucially) reasons for belief, as well as moral reasons. Bart Streumer claims that we are literally unable to believe this broader error theory, for the following reasons. 1. This error theory implies that there is no reason to believe this error theory (there are no reasons at all, so a fortiori there are no reasons to believe this error theory), and anyone who understands it well enough to be in a position to believe it would have to know this. 2. We can't believe something if we believe that there is no reason to believe it. 3. Therefore, we can't believe this error theory. Again, our belief in it would be in a certain way self-defeating.
Third example (Spoilers for Scott Alexander's novel Unsong):
Gur fgbevrf bs gur Pbzrg Xvat naq Ryvfun ora Nohlnu ner nyfb rknzcyrf bs guvf "frys-qrsrngvat ernfbaf" curabzraba. Gur Pbzrg Xvat pna'g tb vagb Uryy orpnhfr va beqre gb tb vagb Uryy ur jbhyq unir gb or rivy. Ohg ur pna'g whfg qb rivy npgf va beqre gb vapernfr uvf "rivy fpber" orpnhfr nal rivy npgf ur qvq jbhyq hygvzngryl or va gur freivpr bs gur tbbq (tbvat vagb Uryy va beqre gb qrfgebl vg), naq gurersber jbhyqa'g pbhag gb znxr uvz rivy. Fb ur pna'g npphzhyngr nal rivy gb trg vagb Uryy gb qrfgebl vg. Ntnva, uvf ernfba sbe tbvat vagb Uryy qrsrngrq uvf novyvgl gb npghnyyl trg vagb Uryy.
What all of these examples have in common is that someone's reason for doing something directly makes it the case that they can't do it. Unless they can find a different reason to do the thing, they won't be able to do it at all.
A couple of meta-notes about shortform content and the frontpage "comments" section:
Roughly agreed (at least with the underlying issues).
We have some vague plans for what I think are strict-improvements over the current status quo (on the front page, making it so you can quickly see all new comments from a given post, and making it so that you can easily see the parents of a comment on the frontpage to get more context), as well as more complex solutions roughly-in-the-direction of what you suggest in the third bullet point (although with different implementation details).
Thinking of running a reading group/crash course on AI alignment-related stuff in my philosophy department. My thought process is that there's been at least some calls for more philosophers to contribute to AI alignment, and there's been some related work done in typical philosophy journals since then (by people like Leonard Dung, Simon Goldstein, etc.) but there is a lot of background that a philosopher new to the field will be missing if they haven't been following e.g. LessWrong for the past decade or so. So I want to construct a crash course in that kind of "AI alignment/LessWrong canon," aimed primarily at philosophers. Goal is to include a mix of generally important/influential stuff, stuff where a philosopher reading it will naturally get excited by some philosophical aspect of it, work by philosophers adjacent to the community (e.g. Joe Carlsmith) to demonstrate that philosophers have a role to play, and direct calls for work by philosophers. Time frame probably 12-ish weeks.
Here is a very rough first draft/longlist. Would love to hear any suggestions!
I said in this comment that I would post an update as to whether or not I had done deep reflection (operationalized as 5 days = 40 hours cumulatively) on AI timelines by August 15th. As it turns out, I have not done so. I had one conversation that caused me to reflect that perhaps timelines are not as key of a variable in my decision process (about whether to drop everything and try to retrain to be useful for AI safety) as I thought they were, but that is the extent of it. I'm not going to commit to do anything further with this right now, because I don't think that would be useful.
I think one of the key distinctions between content that feels "shortform" and content that feels okay to post as a top-level post is that shortform content is content that doesn't feel important/well-developed/long/something enough to have a title. Now, this can't be the whole story, because I have several posts on this very shortform feed that have titles, but it feels like an important piece of the distinction.
With LLMs, as with humans, there are a number of different candidates for "personal identity" over time. In the LLM case some tempting candidates (non-exhaustive) are the model weights, the message history of a thread, the computations underlying a single message or token generation, and the persona (see The Artificial Self and The Persona Selection Model).
But in the LLM case, there seem to be several different dimensions of identity, which can come apart, that don't normally come apart in the human case: specifically consciousness, welfare, self-concern, and predictive utility. Let me briefly unpack:
What this suggests, I think, is that "personal identity" is not actually a natural kind when applied to LLMs. This is a more thoroughgoing skepticism about personal identity than the Parfitian idea that personal identity is vague. Parfit says: there's no further fact about identity; the question has no determinate answer. But as far as I know he's still treating it as one question. And this is plausibly the right move when thinking about the human case. In LLMs, by contrast, not only is there "no further fact" about identity, I suspect there is not even one unified question that can be asked. The different questions that make up identity questions come apart in ways they don't really in humans. (Though if this is right, it arguably implies that even in the human case "personal identity" was always a cluster concept whose components happened to converge.)
As Raemon has suggested a format for a short-form content feed, I'm going to go ahead and make one. His explanation of the format:
I'll be using the comment section of this post as a repository of short, sometimes-half-baked posts that either:
I ask people not to create top-level comments here, but feel free to reply to comments like you would a FB post.
Edit: For me at least, #2 also includes "writing them up as a full post would involve enough effort that the post would probably otherwise not get written."