LESSWRONG
is fundraising!
LW

Richard Ngo's Shortform — LessWrong

Richard Ngo's Shortform

26th Apr 2020

1 min read

6 Ω 3

This is a special post for quick takes by Richard_Ngo. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.

Mentioned in

177Some conceptual alignment research projects

170The behavioral selection model for predicting AI motivations

110How training-gamers might function (and win)

108Value systematization: how values become coherent (and misaligned)

90What goals will AIs have? A list of hypotheses

Load More (5/17)

Richard Ngo's Shortform

5Alexander Gietelink Oldenziel

2plex

2kave

2Richard_Ngo

1Sheikh Abdur Raheem Ali

1Bogdan Ionut Cirstea

4Adrià Garriga-alonso

5the gears to ascension

11the gears to ascension

2the gears to ascension

2Adrià Garriga-alonso

2Richard_Ngo

2Adrià Garriga-alonso

3Richard_Ngo

2Adrià Garriga-alonso

4the gears to ascension

41a3orn

4Daniel Kokotajlo

2the gears to ascension

517 comments, sorted by

top scoring

Click to highlight new comments since: Today at 11:41 AM

Some comments are truncated due to high volume. (⌘F to expand all)Change truncation settings

[-]Richard_Ngo1mo*14241

Here's a list of my donations so far this year (put together as part of thinking through whether I and others should participate in an OpenAI equity donation round).

They are roughly in chronological order (though it's possible I missed one or two). I include some thoughts on what I've learned and what I'm now doing differently at the bottom.

$100k to Lightcone
1. This grant was largely motivated by my respect for Oliver Habryka's quality of thinking and personal judgment.
2. This ended up being matched by the Survival and Flourishing Fund (though I didn't know it would be when I made it). Note that they'll continue matching donations to Lightcone until the end of March 2026.
$50k to the Alignment of Complex Systems (ACS) research group
1. This grant was largely motivated by my respect for Jan Kulveit's philosophical and technical thinking.
$20k to Alexander Gietelink Oldenziel for support with running agent foundations conferences.
~$25k to Inference Magazine to host a public debate on the plausibility of the intelligence explosion in London.
$100k to Apart Research, who run hackathons where people can engage with AI safety research in a hands-on way (technically made with my regranting funds from

... (read more)

[-]maxnadeau1mo124

I'm glad you're betting on your own taste/expertise instead of donating on behalf of the community—that seems sensible to me.

[-]Erich_Grunewald1mo100

What are the best places to start reading about why you are uninterested in almost all commonly proposed AI governance interventions, and about the AI governance interventions you are interested in? I imagine the curriculum sheds some light on this, but it's quite long.

4Richard_Ngo1mo

Hmm, I don't have anything substantive out on this specifically; the closest is probably this talk (though note that some of my arguments in it were a bit sloppy, e.g. as per the top comment).

5Alexander Gietelink Oldenziel1mo

>> " Even more recently, I've decided that I can bet on my research taste most effectively by simply hiring research assistants to work for me. " This seems right to me - plausibly a massive multiplier. For example, John Wentworth told me he found a large productivity boost when he started working David Lorell.

2plex1mo

This seems like both a good process, using your existing knowledge to find good opportunities rather than doing normal applications seems in line with my guess at how high EV grants her name, and a set of grantees I am generally glad to see funded.

2kave1mo

*Lighthaven->Lightcone (at least in the case of SFF matching)

2Richard_Ngo1mo

Ty, fixed.

1Sheikh Abdur Raheem Ali1mo

Thank you for donating!

[-]Richard_Ngo2yΩ7113781

I feel kinda frustrated whenever "shard theory" comes up in a conversation, because it's not a theory, or even a hypothesis. In terms of its literal content, it basically seems to be a reframing of the "default" stance towards neural networks often taken by ML researchers (especially deep learning skeptics), which is "assume they're just a set of heuristics".

This is a particular pity because I think there's a version of the "shard" framing which would actually be useful, but which shard advocates go out of their way to avoid. Specifically: we should be interested in "subagents" which are formed via hierarchical composition of heuristics and/or lower-level subagents, and which are increasingly "goal-directed" as you go up the hierarchy. This is an old idea, FWIW; e.g. it's how Minsky frames intelligence in Society of Mind. And it's also somewhat consistent with the claim made in the original shard theory post, that "shards are just collections of subshards".

The problem is the "just". The post also says "shards are not full subagents", and that "we currently estimate that most shards are 'optimizers' to the extent that a bacterium or a thermostat is an optimizer." But the whole point... (read more)

[-]Daniel Kokotajlo2y161

I am not as negative on it as you are -- it seems an improvement over the 'Bag O' Heuristics' model and the 'expected utility maximizer' model. But I agree with the critique and said something similar here:

you go on to talk about shards eventually values-handshaking with each other. While I agree that shard theory is a big improvement over the models that came before it (which I call rational agent model and bag o' heuristics model) I think shard theory currently has a big hole in the middle that mirrors the hole between bag o' heuristics and rational agents. Namely, shard theory currently basically seems to be saying "At first, you get very simple shards, like the following examples: IF diamond-nearby THEN goto diamond. Then, eventually, you have a bunch of competing shards that are best modelled as rational agents; they have beliefs and desires of their own, and even negotiate with each other!" My response is "but what happens in the middle? Seems super important! Also haven't you just reproduced the problem but inside the head?" (The problem being, when modelling AGI we always understood that it would start out being just a crappy bag of heuristics and end up a scary rational ag

... (read more)

9TurnTrout2y

Personally, I'm not ignoring that question, and I've written about it (once) in some detail. Less relatedly, I've talked about possible utility function convergence via e.g. A shot at the diamond-alignment problem and my recent comment thread with Wei_Dai. It's not that there isn't more shard theory content which I could write, it's that I got stuck and burned out before I could get past the 101-level content. I felt * a) gaslit by "I think everyone already knew this" or even "I already invented this a long time ago" (by people who didn't seem to understand it); and that * b) I wasn't successfully communicating many intuitions;[1] and * c) it didn't seem as important to make theoretical progress anymore, especially since I hadn't even empirically confirmed some of my basic suspicions that real-world systems develop multiple situational shards (as I later found evidence for in Understanding and controlling a maze-solving policy network). So I didn't want to post much on the site anymore because I was sick of it, and decided to just get results empirically. I've always read "assume heuristics" as expecting more of an "ensemble of shallow statistical functions" than "a bunch of interchaining and interlocking heuristics from which intelligence is gradually constructed." Note that (at least in my head) the shard view is extremely focused on how intelligence (including agency) is comprised of smaller shards, and the developmental trajectory over which those shards formed. 1. ^ The 2022 review indicates that more people appreciated the shard theory posts than I realized at the time.

[-]Daniel Kokotajlo2yΩ8100

FWIW I'm potentially intrested in interviewing you (and anyone else you'd recommend) and then taking a shot at writing the 101-level content myself.

2Daniel Kokotajlo2y

Curious to hear whether I was one of the people who contributed to this.

4TurnTrout2y

Nope! I have basically always enjoyed talking with you, even when we disagree.

2Daniel Kokotajlo2y

Ok, whew, glad to hear.

4tailcalled2y

But shard theorists mainly aim to address agency obtained via DPO-like setups, and @TurnTrout has mathematically proved that such setups don't favor the power-seeking drives AI safety researchers are usually concerned about in the context of agency.

2rotatingpaguro2y

I read the section you linked, but I can't follow it. Anyway, here it is its conclusive paragraph: From this alone, I get the impression that he hasn't proved that "there isn't instrumental convergence", but that "there isn't a totally general instrumental convergence that applies even to very wild utility functions".

3tailcalled2y

A key part of instrumental convergence is the convergence aspect, which as I understand it refers to the notion that even very wild utility functions will share certain preferences. E.g. the empirical tendency for random chess board evaluations to prefer mobility. If you don't have convergence, you don't have instrumental convergence.

3rotatingpaguro2y

Ok. Then I'll say that randomly assigned utility over full trajectories are beyond wild! The basin of attraction just needs to be large enough. AIs will intentionally be created with more structure than that.

5tailcalled2y

The issue isn't the "full trajectories" part; that actually makes instrumental convergence stronger. The issue is the "actions" part. In terms of RLHF, what this means is that people might not simply blindly follow the instructions given by AIs and rate them based on the ultimate outcome (even if the outcome differs wildly from what they'd intuitively think it'd do), but rather they might think about the instructions the AIs provide, and rate them based on whether they a priori make sense. If the AI then has some galaxybrained method of achieving something (which traditionally would be instrumentally convergent) that humans don't understand, then that method will be negatively reinforced (because people don't see the point of it and therefore downvote it), which eliminates dangerous powerseeking.

[-]Richard_Ngo9mo*9321

In response to an email about what a pro-human ideology for the future looks like, I wrote up the following:

The pro-human egregore I'm currently designing (which I call fractal empowerment) incorporates three key ideas:

Firstly, we can see virtue ethics as a way for less powerful agents to aggregate to form more powerful superagents that preserve the interests of those original less powerful agents. E.g. virtues like integrity, loyalty, etc help prevent divide-and-conquer strategies. This would have been in the interests of the rest of the world when Europe was trying to colonize them, and will be in the best interests of humans when AIs try to conquer us.

Secondly, the most robust way for a more powerful agent to be altruistic towards a less powerful agent is not for it to optimize for that agent's welfare, but rather to optimize for its empowerment. This prevents predatory strategies from masquerading as altruism (e.g. agents claiming "I'll conquer you and then I'll empower you", which then somehow never get around to the second step).

Thirdly: the generational contract. From any given starting point, there are a huge number of possible coalitions which could form, an... (read more)

[-]Wei Dai9mo*3810

How would this ideology address value drift? I've been thinking a lot about the kind quoted in Morality is Scary. The way I would describe it now is that human morality is by default driven by a competitive status/signaling game, where often some random or historically contingent aspect of human value or motivation becomes the focal point of the game, and gets magnified/upweighted as a result of competitive dynamics, sometimes to an extreme, even absurd degree.

(Of course from the inside it doesn't look absurd, but instead feels like moral progress. One example of this that I happened across recently is filial piety in China, which became more and more extreme over time, until someone cutting off a piece of their flesh to prepare a medicinal broth for an ailing parent was held up as a moral exemplar.)

Related to this is my realization is that the kind of philosophy you and I are familiar with (analytical philosophy, or more broadly careful/skeptical philosophy) doesn't exist in most of the world and may only exist in Anglophone countries as a historical accident. There, about 10,000 practitioners exist who are funded but ignored by the rest of the population. To most of humanity, "ph... (read more)

2Richard_Ngo9mo

One of the main ways I think about empowerment is in terms of allowing better coordination between subagents. In the case of an individual human, extreme morality can be seen as one subagent seizing control and overriding other subagents (like the ones who don't want to chop off body parts). In the case of a group, extreme morality can be seen in terms of preference cascades that go beyond what most (or even any) of the individuals involved with them would individually prefer. In both cases, replacing fear-based motivation with less coercive/more cooperative interactions between subagents would go a long way towards reducing value drift.

7Wei Dai9mo

I'm not sure that fear or coercion has much to do with it, because there's often no internal conflict when someone is caught up in some extreme form of the morality game, they're just going along with it wholeheartedly, thinking they're just being a good person or helping to advance the arc of history. In the subagents frame, I would say that the subagents have an implicit contract/agreement that any one of them can seize control, if doing so seems good for the overall agent in terms of power or social status. But quite possibly I'm not getting your point, in which case please explain more, or point to some specific parts of your articles that are especially relevant?

4Richard_Ngo4mo

Belated reply, sorry, but I basically just think that this is false—analogous to a dictator who cites parades where people are forced to attend and cheer as evidence that his country lacks internal conflict. Instead, the internal conflict has just been rendered less legible. Note that this is an extremely non-robust agent design! In particular, it allows subagents to gain arbitrary amounts of power simply by lying about their intentions. If you encounter an agent which considers itself to be structured like this, you should have a strong prior that it is deceiving itself about the presence of more subtle control mechanisms.

[-]quetzal_rainbow9mo142

In some sense this is a core idea of UDT: when coordinating with forks of yourself, you defer to your unique last common ancestor. When it's not literally a fork of yourself, there's more arbitrariness but you can still often find a way to use history to narrow down on coordination Schelling points (e.g. "what would Jesus do")

I think this is wholly incorrect line of thinking. UDT operates on your logical ancestor, not literal.

Say, if you know enough science, you know that normal distribution is a maxentropy distribution for fixed mean and variance, and therefore, optimal prior distribution under certain set of assumptions. You can ask yourself question "let's suppose that I haven't seen this evidence, what would be my prior probability?" and get an answer and cooperate with your counterfactual versions which have seen other versions of evidence. But you can't cooperate with your hypothetical version which doesn't know what normal distribution is, because, if it doesn't know about normal distribution, it can't predict how you would behave and account for this in cooperation.

Sufficiently different versions of yourself are just logically uncorrelated with you and there is no game-theoretic reason to account for them.

4Richard_Ngo4mo

Seems odd to make an absolute statement here. More different versions of yourself are less and less correlated, but there's still some correlation. And UDT should also be applicable to interactions with other people, who are typically different from you in a whole bunch of ways.

2quetzal_rainbow4mo

Absolute sense comes from absolute nature of taking actions, not absolute nature of logical correlation. I.e., in Prisoner's Dilemma with payoffs (5,5)(10,1)(2,2) you should defect if your counterparty is capable to act conditional on your action in less than 75% of cases, which is quite high logical correlation, but expected value is higher if you defect.

5Purplehermann9mo

2nd point is a scary one. Empowering others in the relative sense is a terrible idea, unless they are trustworthy/virtuous. Same issue as AI risk In the absolute terms sure

2Nathan Helm-Burger9mo

Yeah, this definitely needs to be a limited sorry of empowerment, in my mind. Like, imagine you wanted to give a 5 year old child the best day ever. You wanted to give them really fun options, but also not cause them to suffer from decision fatigue, or regret about the paths not taken. More importantly, if they asked for an alien ray gun with which to shoot bad guys, giving them an actual extremely dangerous weapon would be a terrible idea. Similarly, offering them a ride on a cool looking roller coaster that was actually a 'death coaster' would be a terrible trap.

4Noosphere898mo

(I'm inspired to write this comment on the notion of empowerment because of Richard Ngo's recent comment in another post on Towards a scale-free theory of intelligent agency, so I'll both respond to the empowerment motion and part of the comment linked below: https://www.lesswrong.com/posts/5tYTKX4pNpiG4vzYg/towards-a-scale-free-theory-of-intelligent-agency#nigkBt47pLMi5tnGd): To address this part of the linked comment: This depends a lot on how the conflict started, and in particular, I don't think that we should do something else if the conflict arose out of severe alignment failures of AIs/radically augmented people, since the generational contract/UDT/LDT/FDT cannot be used as a substitute for alignment (this was the takeaway I got from reading Nate Soares's post on Decision theory does not imply that we get to have nice things, and while @ryan_greenblatt commented that it's unlikely that alignment failures end up in us getting extinct without decision theory saving us, note that us being saved can still be really, really rough, and probably ends up with billions of present humans dying without everyone else dying (though I don't think about decision theory much), so conflicts with future agents cannot always be avoided if we mess up hard enough on the alignment problems of the future). https://www.lesswrong.com/posts/rP66bz34crvDudzcJ/decision-theory-does-not-imply-that-we-get-to-have-nice Now I'll address the fractal empowerment idea. I like the idea of fractal empowerment, and IMO is one of my biggest ideals if we succeed at getting out of the risky state we are in, only rivaled by infra-Bayes Physicalism's plan for alignment, which I currently think is called Physical Super-Imitation after the monotonicity principle managed to be removed, meaning way more preferences could be fit into than before: https://www.lesswrong.com/posts/DobZ62XMdiPigii9H/non-monotonic-infra-bayesian-physicalism https://www.lesswrong.com/posts/ZwshvqiqCvXPsZEct/the-learning-t

2William_S9mo

Would be interested in a quick write-up of what you think are the most important virtues you'd want for AI systems, seems good in terms of having things to aim towards instead of just aiming away from.

2Fabien Roger9mo

I did not understand that. Is the worry that it's hard to distinguish a genuine welfare maximizer from a predator because you can't tell if they will ever give you back power? I don't understand why this does not apply to agents pretending to pursue empowerment. It is common in conflicts to temporarily disempower someone to protect their long-term empowerment (e.g. a country mandatorily mobilizing for war against a fascist attacker, preventing a child from ignoring their homework), and it is also common to pretend to protect long-term empowerment and never give back power (e.g. a dictatorship of the proletariat never transitioning to a "true" communist economy).

4Richard_Ngo9mo

Ah, yeah, I was a bit unclear here. Basic idea is that by conquering someone you may not reduce their welfare very much short-term, but you do reduce their power a lot short-term. (E.g. the British conquered India with relatively little welfare impacts on most Indians.) And so it is much harder to defend a conquest as altruistic in the sense of empowering, than it is to defend a conquest as altruistic in the sense of welfare-increasing. As you say, this is not a perfect defense mechanism, because sometimes long-term empowerment and short-term empowerment conflict. But there are often strategies that are less disempowering short-term which the moral pressure of "altruism=empowerment" would push people towards. E.g. it would make it harder for people to set up the "dictatorship of the proletariat" in the first place. And in general I think it's actually pretty uncommon for temporary disempowerment to be necessary for long-term empowerment. Re your homework example, there's a wide spectrum from the highly-empowering Taking Children Seriously to highly-disempowering Asian tiger parents, and I don't think it's a coincidence that tiger parenting often backfires. Similarly, mandatory mobilization disproportionately happens during wars fought for the wrong reasons.

[-]Richard_Ngo5y768

One fairly strong belief of mine is that Less Wrong's epistemic standards are not high enough to make solid intellectual progress here. So far my best effort to make that argument has been in the comment thread starting here. Looking back at that thread, I just noticed that a couple of those comments have been downvoted to negative karma. I don't think any of my comments have ever hit negative karma before; I find it particularly sad that the one time it happens is when I'm trying to explain why I think this community is failing at its key goal of cultivating better epistemics.

There's all sorts of arguments to be made here, which I don't have time to lay out in detail. But just step back for a moment. Tens or hundreds of thousands of academics are trying to figure out how the world works, spending their careers putting immense effort into reading and producing and reviewing papers. Even then, there's a massive replication crisis. And we're trying to produce reliable answers to much harder questions by, what, writing better blog posts, and hoping that a few of the best ideas stick? This is not what a desperate effort to find the truth looks like.

[-]Trinley Goldenberg5y*13-1

And we're trying to produce reliable answers to much harder questions by, what, writing better blog posts, and hoping that a few of the best ideas stick? This is not what a desperate effort to find the truth looks like.

It seems to me that maybe this is what a certain stage in the desperate effort to find the truth looks like?

Like, the early stages of intellectual progress look a lot like thinking about different ideas and seeing which ones stand up robustly to scrutiny. Then the best ones can be tested more rigorously and their edges refined through experimentation.

It seems to me like there needs to be some point in the desparate search for truth in which you're allowing for half-formed thoughts and unrefined hypotheses, or else you simply never get to a place where the hypotheses you're creating even brush up against the truth.

[-]Richard_Ngo5y150

In the half-formed thoughts stage, I'd expect to see a lot of literature reviews, agendas laying out problems, and attempts to identify and question fundamental assumptions. I expect that (not blog-post-sized speculation) to be the hard part of the early stages of intellectual progress, and I don't see it right now.

Perhaps we can split this into technical AI safety and everything else. Above I'm mostly speaking about "everything else" that Less Wrong wants to solve. Since AI safety is now a substantial enough field that its problems need to be solved in more systemic ways.

8Trinley Goldenberg5y

I would expect that later in the process. Agendas laying out problems and fundamental assumptions don't spring from nowhere (at least for me), they come from conversations where I'm trying to articulate some intuition, and I recognize some underlying pattern. The pattern and structure doesn't emerge spontaneously, it comes from trying to pick around the edges of a thing, get thoughts across, explain my intuitions and see where they break. I think it's fair to say that crystallizing these patterns into a formal theory is a "hard part", but the foundation for making it easy is laid out in the floundering and flailing that came before.

[-]Past Account5y*100

[Deleted]

3Viliam5y

Ironically, some people already feel threatened by the high standards here. Setting them higher probably wouldn't result in more good content. It would result in less mediocre content, but probably also less good content, as the authors who sometimes write a mediocre article and sometimes a good one, would get discouraged and give up. Ben Pace gives a few examples of great content in the next comment. It would be better to easier separate the good content from the rest, but that's what the reviews are for. Well, only one review so far, if I remember correctly. I would love to see reviews of pre-2018 content (maybe multiple years in one review, if they were less productive). Then I would love to see the winning content get the same treatment as the Sequences -- edit them and arrange them into a book, and make it "required reading" for the community (available as a free PDF).

7Past Account5y

[Deleted]

6Ben Pace5y

The top posts in the 2018 Review are filled with fascinating and well-explained ideas. Many of the new ideas are not settled science, but they're quite original and substantive, or excellent distillations of settled science, and are often the best piece of writing on the internet about their topics. You're wrong about LW epistemic standards not being high enough to make solid intellectual progress, we already have. On AI alone (which I am using in large part because there's vaguely more consensus around it than around rationality), I think you wouldn't have seen almost any of the public write-ups (like Embedded Agency and Zhukeepa's Paul FAQ) without LessWrong, and I think a lot of them are brilliant. I'm not saying we can't do far better, or that we're sufficiently good. Many of the examples of success so far are "Things that were in people's heads but didn't have a natural audience to share them with". There's not a lot of collaboration at present, which is why I'm very keen to build the new LessWrong Docs that allows for better draft sharing and inline comments and more. We're working on the tools for editing tags, things like edit histories and so on, that will allow us to build a functioning wiki system to have canonical writeups and explanation that people add to and refine. I want future iterations of the LW Review to have more allowance for incorporating feedback from reviewers. There's lots of work to do, and we're just getting started. But I disagree the direction isn't "a desperate effort to find the truth". That's what I'm here for. Even in the last month or two, how do you look at things like this and this and this and this and not think that they're likely the best publicly available pieces of writing in the world about their subjects? Wrt rationality, I expect things like this and this and this and this will probably go down as historically important LW posts that helped us understand the world, and make a strong showing in the 2020 LW Review.

[-]Richard_Ngo5y*420

As mentioned in my reply to Ruby, this is not a critique of the LW team, but of the LW mentality. And I should have phrased my point more carefully - "epistemic standards are too low to make any progress" is clearly too strong a claim, it's more like "epistemic standards are low enough that they're an important bottleneck to progress". But I do think there's a substantive disagreement here. Perhaps the best way to spell it out is to look at the posts you linked and see why I'm less excited about them than you are.

Of the top posts in the 2018 review, and the ones you linked (excluding AI), I'd categorise them as follows:

Interesting speculation about psychology and society, where I have no way of knowing if it's true:

Local Validity as a Key to Sanity and Civilization
The Loudest Alarm Is Probably False
Anti-social punishment (which is, unlike the others, at least based on one (1) study).
Babble
Intelligent social web
Unrolling social metacognition
Simulacra levels
Can you keep this secret?

Same as above but it's by Scott so it's a bit more rigorous and much more compelling:

Is Science Slowing Down?
The tails coming apart as a metaph

... (read more)

[-]Ruby5y120

(Thanks for laying out your position in this level of depth. Sorry for how long this comment turned out. I guess I wanted to back up a bunch of my agreement with words. It's a comment for the sake of everyone else, not just you.)

I think there's something to what you're saying, that the mentality itself could be better. The Sequences have been criticized because Eliezer didn't cite previous thinkers all that much, but at least as far as the science goes, as you said, he was drawing on academic knowledge. I also think we've lost something precious with the absence of epic topic reviews by the likes of Luke. Kaj Sotala still brings in heavily from outside knowledge, John Wentworth did a great review on Biological Circuits, and we get SSC crossposts that have that, but otherwise posts aren't heavily referencing or building upon outside stuff. I concede that I would like to see a lot more of that.

I think Kaj was rightly disappointed that he didn't get more engagement with his post whose gist was "this is what the science really says about S1 & S2, one of your most cherished concepts, LW community".

I wouldn't say the typical approach is strictly bad, there's value in thinking freshly... (read more)

[-]drossbucket5y410

This is only tangentially relevant, but adding it here as some of you might find it interesting:

Venkatesh Rao has an excellent Twitter thread on why most independent research only reaches this kind of initial exploratory level (he tried it for a bit before moving to consulting). It's pretty pessimistic, but there is a somewhat more optimistic follow-up thread on potential new funding models. Key point is that the later stages are just really effortful and time-consuming, in a way that keeps out a lot of people trying to do this as a side project alongside a separate main job (which I think is the case for a lot of LW contributors?)

Quote from that thread:

Research =

a) long time between having an idea and having something to show for it that even the most sympathetic fellow crackpot would appreciate (not even pay for, just get)

b) a >10:1 ratio of background invisible thinking in notes, dead-ends, eliminating options etc

With a blogpost, it’s like a week of effort at most from idea to mvp, and at most a 3:1 ratio of invisible to visible. That’s sustainable as a hobby/side thing.

To do research-grade thinking you basically have to be independently wealthy and accept 90% d

... (read more)

[-]Richard_Ngo5y140

Thanks, these links seem great! I think this is a good (if slightly harsh) way of making a similar point to mine:

"I find that autodidacts who haven’t experienced institutional R&D environments have a self-congratulatory low threshold for what they count as research. It’s a bit like vanity publishing or fan fiction. This mismatch doesn’t exist as much in indie art, consulting, game dev etc"

6Richard_Ngo5y

Also, I liked your blog post! More generally, I strongly encourage bloggers to have a "best of" page, or something that directs people to good posts. I'd be keen to read more of your posts but have no idea where to start.

6drossbucket5y

Thanks! I have been meaning to add a 'start here' page for a while, so that's good to have the extra push :) Seems particularly worthwhile in my case because a) there's no one clear theme and b) I've been trying a lot of low-quality experimental posts this year bc pandemic trashed motivation, so recent posts are not really reflective of my normal output. For now some of my better posts in the last couple of years might be Cognitive decoupling and banana phones (tracing back the original precursor of Stanovich's idea), The middle distance (a writeup of a useful and somewhat obscure idea from Brian Cantwell Smith's On the Origin of Objects), and the negative probability post and its followup.

[-]Ben Pace5y100

Quoting your reply to Ruby below, I agree I'd like LessWrong to be much better at "being able to reliably produce and build on good ideas".

The reliability and focus feels most lacking to me on the building side, rather than the production, which I think we're doing quite well at. I think we've successfully formed a publishing platform that provides and audience who are intensely interested in good ideas around rationality, AI, and related subjects, and a lot of very generative and thoughtful people are writing down their ideas here.

We're low on the ability to connect people up to do more extensive work on these ideas – most good hypotheses and arguments don't get a great deal of follow up or further discussion.

Here are some subjects where I think there's been various people sharing substantive perspectives, but I think there's also a lot of space for more 'details' to get fleshed out and subquestions to be cleanly answered:

Sabbath and Rest Days (Zvi, Lauren Lee, Jacobian, Scott)
Moloch and Slack and Mazes (Scott, Eliezer, Zvi, Swentworth, Jameson)
Inner/Outer Alignment (EvHub, Rafael, Paul, Swentworth, Steve2152)
Embedded Agency + Optimization (Abram, Scott, Swentworth, Alex Fli

... (read more)

[-]Richard_Ngo5y120

"I see a lot of (very high quality) raw energy here that wants shaping and directing, with the use of lots of tools for coordination (e.g. better collaboration tools)."

Yepp, I agree with this. I guess our main disagreement is whether the "low epistemic standards" framing is a useful way to shape that energy. I think it is because it'll push people towards realising how little evidence they actually have for many plausible-seeming hypotheses on this website. One proven claim is worth a dozen compelling hypotheses, but LW to a first approximation only produces the latter.

When you say "there's also a lot of space for more 'details' to get fleshed out and subquestions to be cleanly answered", I find myself expecting that this will involve people who believe the hypothesis continuing to build their castle in the sky, not analysis about why it might be wrong and why it's not.

That being said, LW is very good at producing "fake frameworks". So I don't want to discourage this too much. I'm just arguing that this is a different thing from building robust knowledge about the world.

6Ben Pace5y

I will continue to be contrary and say I'm not sure I agree with this. For one, I think in many domains new ideas are really hard to come by, as opposed to making minor progress in the existing paradigms. Fundamental theories in physics, a bunch of general insights about intelligence (in neuroscience and AI), etc. And secondly, I am reminded of what Lukeprog wrote in his moral consciousness report, that he wished the various different philosophies-of-consciousness would stop debating each other, go away for a few decades, then come back with falsifiable predictions. I sometimes take this stance regarding many disagreements of import, such as the basic science vs engineering approaches to AI alignment. It's not obvious to me that the correct next move is for e.g. Eliezer and Paul to debate for 1000 hours, but instead to go away and work on their ideas for a decade then come back with lots of fleshed out details and results that can be more meaningfully debated. I feel similarly about simulacra levels, Embedded Agency, and a bunch of IFS stuff. I would like to see more experimentation and literature reviews where they make sense, but I also feel like these are implicitly making substantive and interesting claims about the world, and I'd just be interested in getting a better sense of what claims they're making, and have them fleshed out + operationalized more. That would be a lot of progress to me, and I think each of them is seeing that sort of work (with Zvi, Abram, and Kaj respectively leading the charges on LW, alongside many others).

[-]Richard_Ngo5y160

I feel like this comment isn't critiquing a position I actually hold. For example, I don't believe that "the correct next move is for e.g. Eliezer and Paul to debate for 1000 hours". I am happy for people to work towards building evidence for their hypotheses in many ways, including fleshing out details, engaging with existing literature, experimentation, and operationalisation.

Perhaps this makes "proven claim" a misleading phrase to use. Perhaps more accurate to say: "one fully fleshed out theory is more valuable than a dozen intuitively compelling ideas". But having said that, I doubt that it's possible to fully flesh out a theory like simulacra levels without engaging with a bunch of academic literature and then making predictions.

I also agree with Raemon's response below.

[-]Raemon5y110

I think I'm concretely worried that some of those models / paradigms (and some other ones on LW) don't seem pointed in a direction that leads obviously to "make falsifiable predictions."

And I can imagine worlds where "make falsifiable predictions" isn't the right next step, you need to play around with it more and get it fleshed out in your head before you can do that. But there is at least some writing on LW that feels to me like it leaps from "come up with an interesting idea" to "try to persuade people it's correct" without enough checking.

(In the case of IFS, I think Kaj's sequence is doing a great job of laying it out in a concrete way where it can then be meaningfully disagreed with. But the other people who've been playing around with IFS didn't really seem interested in that, and I feel like we got lucky that Kaj had the time and interest to do so.)

4Ben Pace5y

A housemate of mine said to me they think LW has a lot of breadth, but could benefit from more depth. I think in general when we do intellectual work we have excellent epistemic standards, capable of listening to all sorts of evidence that other communities and fields would throw out, and listening to subtler evidence than most scientists ("faster than science"), but that our level of coordination and depth is often low. "LessWrongers should collaborate more and go into more depth in fleshing out their ideas" sounds more true to me than "LessWrongers have very low epistemic standards".

[-]Richard_Ngo5y200

In general when we do intellectual work we have excellent epistemic standards, capable of listening to all sorts of evidence that other communities and fields would throw out, and listening to subtler evidence than most scientists ("faster than science")

"Being more openminded about what evidence to listen to" seems like a way in which we have lower epistemic standards than scientists, and also that's beneficial. It doesn't rebut my claim that there are some ways in which we have lower epistemic standards than many academic communities, and that's harmful.

In particular, the relevant question for me is: why doesn't LW have more depth? Sure, more depth requires more work, but on the timeframe of several years, and hundreds or thousands of contributors, it seems viable. And I'm proposing, as a hypothesis, that LW doesn't have enough depth because people don't care enough about depth - they're willing to accept ideas even before they've been explored in depth. If this explanation is correct, then it seems accurate to call it a problem with our epistemic standards - specifically, the standard of requiring (and rewarding) deep investigation and scholarship.

7John_Maxwell5y

Your solution to the "willingness to accept ideas even before they've been explored in depth" problem is to explore ideas in more depth. But another solution is to accept fewer ideas, or hold them much more provisionally. I'm a proponent of the second approach because: * I suspect even academia doesn't hold ideas as provisionally as it should. See Hamming on expertise: https://forum.effectivealtruism.org/posts/mG6mckPHAisEbtKv5/should-you-familiarize-yourself-with-the-literature-before?commentId=SaXXQXLfQBwJc9ZaK * I suspect trying to browbeat people to explore ideas in more depth works against the grain of an online forum as an institution. Browbeating works in academia because your career is at stake, but in an online forum, it just hurts intrinsic motivation and cuts down on forum use (the forum runs on what Clay Shirky called "cognitive surplus", essentially a term for peoples' spare time and motivation). I'd say one big problem with LW 1.0 that LW 2.0 had to solve before flourishing was people felt too browbeaten to post much of anything. If we accept fewer ideas / hold them much more provisionally, but provide a clear path to having an idea be widely held as true, that creates an incentive for people to try & jump through hoops--and this incentive is a positive one, not a punishment-driven browbeating incentive. Maybe part of the issue is that on LW, peer review generally happens in the comments after you publish, not before. So there's no publication carrot to offer in exchange for overcoming the objections of peer reviewers.

4Richard_Ngo5y

"If we accept fewer ideas / hold them much more provisionally, but provide a clear path to having an idea be widely held as true, that creates an incentive for people to try & jump through hoops--and this incentive is a positive one, not a punishment-driven browbeating incentive." Hmm, it sounds like we agree on the solution but are emphasising different parts of it. For me, the question is: who's this "we" that should accept fewer ideas? It's the set of people who agree with my argument that you shouldn't believe things which haven't been fleshed out very much. But the easiest way to add people to that set is just to make the argument, which is what I've done. Specifically, note that I'm not criticising anyone for producing posts that are short and speculative: I'm criticising the people who update too much on those posts.

4John_Maxwell5y

Fair enough. I'm reminded of a time someone summarized one of my posts as being a definitive argument against some idea X and me thinking to myself "even I don't think my post definitively settles this issue" haha.

3Raemon5y

Yeah, this is roughly how I think about it. I do think right now LessWrong should lean more in the direction the Richard is suggesting – I think it was essential to establish better Babble procedures but now we're doing well enough on that front that I think setting clearer expectations of how the eventual pruning works is reasonable.

7Richard_Ngo5y

I wanted to register that I don't like "babble and prune" as a model of intellectual development. I think intellectual development actually looks more like: 1. Babble 2. Prune 3. Extensive scholarship 4. More pruning 5. Distilling scholarship to form common knowledge And that my main criticism is the lack of 3 and 5, not the lack of 2 or 4. I also note that: a) these steps get monotonically harder, so that focusing on the first two misses *almost all* the work; b) maybe I'm being too harsh on the babble and prune framework because it's so thematically appropriate for me to dunk on it here; I'm not sure if your use of the terminology actually reveals a substantive disagreement.

2Raemon5y

I basically agree with your 5-step model (I at least agree it's a more accurate description than Babel and Prune, which I just meant as rough shorthand). I'd add things like "original research/empiricism" or "more rigorous theorizing" to the "Extensive Scholarship" step. I see the LW Review as basically the first of (what I agree should essentially be at least) a 5 step process. It's adding a stronger Step 2, and a bit of Step 5 (at least some people chose to rewrite their posts to be clearer and respond to criticism) ... Currently, we do get non-zero Extensive Scholarship and Original Empiricism. (Kaj's Multi-Agent Models of Mind seems like it includes real scholarship. Scott Alexander / Eli Tyre and Bucky's exploration into Birth Order Effects seemed like real empiricism). Not nearly as much as I'd like. But John's comment elsethread seems significant: This reminded of a couple posts in the 2018 Review, Local Validity as Key to Sanity and Civilization, and Is Clickbait Destroying Our General Intelligence?. Both of those seemed like "sure, interesting hypothesis. Is it real tho?" During the Review I created a followup "How would we check if Mathematicians are Generally More Law Abiding?" question, trying to move the question from Stage 2 to 3. I didn't get much serious response, probably because, well, it was a much harder question. But, honestly... I'm not sure it's actually a question that was worth asking. I'd like to know if Eliezer's hypothesis about mathematicians is true, but I'm not sure it ranks near the top of questions I'd want people to put serious effort into answering. I do want LessWrong to be able to followup Good Hypotheses with Actual Research, but it's not obvious which questions are worth answering. OpenPhil et al are paying for some types of answers, I think usually by hiring researchers full time. It's not quite clear what the right role for LW to play in the ecosystem.

2John_Maxwell5y

1. All else equal, the harder something is, the less we should do it. 2. My quick take is that writing lit reviews/textbooks is a comparative disadvantage of LW relative to the mainstream academic establishment. In terms of producing reliable knowledge... if people actually care about whether something is true, they can always offer a cash prize for the best counterargument (which could of course constitute citation of academic research). The fact that people aren't doing this suggests to me that for most claims on LW, there isn't any (reasonably rich) person who cares deeply re: whether the claim is true. I'm a little wary of putting a lot of effort into supply if there is an absence of demand. (I guess the counterargument is that accurate knowledge is a public good so an individual's willingness to pay doesn't get you complete picture of the value accurate knowledge brings. Maybe what we need is a way to crowdfund bounties for the best argument related to something.) (I agree that LW authors would ideally engage more with each other and academic literature on the margin.)

4DirectedEvolution5y

I’ve been thinking about the idea of “social rationality” lately, and this is related. We do so much here in the way of training individual rationality - the inputs, functions, and outputs of a single human mind. But if truth is a product, then getting human minds well-coordinated to produce it might be much more important than training them to be individually stronger. Just as assembly line production is much more effective in producing almost anything than teaching each worker to be faster in assembling a complete product by themselves. My guess is that this could be effective not only in producing useful products, but also in overcoming biases. Imagine you took 5 separate LWers and asked them to create a unified consensus response to a given article. My guess is that they’d learn more through that collective effort, and produce a more useful response, than if they spent the same amount of time individually evaluating the article and posting their separate replies. Of course, one of the reasons we don’t to that so much is that coordination is an up-front investment and is unfamiliar. Figuring out social technology to make it easier to participate in might be a great project for LW.

[-]John_Maxwell5y110

There's been a fair amount of discussion of that sort of thing here: https://www.lesswrong.com/tag/group-rationality There are also groups outside LW thinking about social technology such as RadicalxChange.

Imagine you took 5 separate LWers and asked them to create a unified consensus response to a given article. My guess is that they’d learn more through that collective effort, and produce a more useful response, than if they spent the same amount of time individually evaluating the article and posting their separate replies.

I'm not sure. If you put those 5 LWers together, I think there's a good chance that the highest status person speaks first and then the others anchor on what they say and then it effectively ends up being like a group project for school with the highest status person in charge. Some related links.

3DirectedEvolution5y

That’s definitely a concern too! I imagine such groups forming among people who either already share a basic common view, and collaborate to investigate more deeply. That way, any status-anchoring effects are mitigated. Alternatively, it could be an adversarial collaboration. For me personally, some of the SSC essays in this format have led me to change my mind in a lasting way.

4Elliot Temple5y

People also reject ideas before they've been explored in depth. I've tried to discuss similar issues with LW before but the basic response was roughly "we like chaos where no one pays attention to whether an argument has ever been answered by anyone; we all just do our own thing with no attempt at comprehensiveness or organizing who does what; having organized leadership of any sort, or anyone who is responsible for anything, would be irrational" (plus some suggestions that I'm low social status and that therefore I personally deserve to be ignored. there were also suggestions – phrased rather differently but amounting to this – that LW will listen more if published ideas are rewritten, not to improve on any flaws, but so that the new versions can be published at LW before anywhere else, because the LW community's attention allocation is highly biased towards that).

2Ben Pace5y

I feel somewhat inclined to wrap up this thread at some point, even while there's more to say. We can continue if you like and have something specific or strong you'd like to ask, but otherwise will pause here.

1TAG5y

You have to realise that what you are doing isn't adequate in order to gain the motivation to do it better, and that is unlikely to happen if you are mostly communicating with other people who think everything is OK.

3TAG5y

Lesswrong is competing against philosophy as well as science, and philosophy has broader criterion of evidence still. In fact , lesswrongians are often frustrated that mainstream philosophy takes such topics as dualism or theism seriously.. even though theres an abundance of Bayesian evidence for them.

2John_Maxwell5y

Depends on the claim, right? If the cost of evaluating a hypothesis is high, and hypotheses are cheap to generate, I would like to generate a great deal before selecting one to evaluate.

2DanielFilan5y

As mentioned in this comment, the Unrolling social metacognition paper is closely related to at least one research paper.

6Richard_Ngo5y

Right, but this isn't mentioned in the post? Which seems odd. Maybe that's actually another example of the "LW mentality": why is the fact that there has been solid empirical research into 3 layers not being enough not important enough to mention in a post on why 3 layers isn't enough? (Maybe because the post was time-boxed? If so that seems reasonable, but then I would hope that people comment saying "Here's a very relevant paper, why didn't you cite it?")

7Past Account5y

[Deleted]

[-]Ben Pace5y110

Much of the same is true of scientific journals. Creating a place to share and publish research is a pretty key piece of intellectual infrastructure, especially for researchers to create artifacts of their thinking along the way.

The point about being 'cross-posted' is where I disagree the most.

This is largely original content that counterfactually wouldn't have been published, or occasionally would have been published but to a much smaller audience. What Failure Looks Like wasn't crossposted, Anna's piece on reality-revealing puzzles wasn't crossposted. I think that Zvi would have still written some on mazes and simulacra, but I imagine he writes substantially more content given the cross-posting available for the LW audience. Could perhaps check his blogging frequency over the last few years to see if that tracks. I recall Zhu telling me he wrote his FAQ because LW offered an audience for it, and likely wouldn't have done so otherwise. I love everything Abram writes, and while he did have the Intelligent Agent Foundations Forum, it had a much more concise, technical style, tiny audience, and didn't have the conversational explanations and stories and cartoons that have... (read more)

5Rohin Shah5y

Yeah, that's true, though it might have happened at some later point in the future as I got increasingly frustrated by people continuing to cite VNM at me (though probably it would have been a blog post and not a full sequence). Reading through this comment tree, I feel like there's a distinction to be made between "LW / AIAF as a platform that aggregates readership and provides better incentives for blogging", and "the intellectual progress caused by posts on LW / AIAF". The former seems like a clear and large positive of LW / AIAF, which I think Richard would agree with. For the latter, I tend to agree with Richard, though perhaps not as strongly as he does. Maybe I'd put it as, I only really expect intellectual progress from a few people who work on problems full time who probably would have done similar-ish work if not for LW / AIAF (but likely would not have made it public). I'd say this mostly for the AI posts. I do read the rationality posts and don't get a different impression from them, but I also don't think enough about them to be confident in my opinions there.

3Ben Pace5y

By "AN" do you mean the AI Alignment Forum, or "AIAF"?

1Past Account5y

[Deleted]

2Ben Pace5y

I did suspect you'd confused it with the Alignment Newsletter :)

5Ruby5y

Thanks for chiming in with this. People criticizing the epistemics is hopefully how we get better epistemics. When the Californian smoke isn't interfering with my cognition as much, I'll try to give your feedback (and Rohin's) proper attention. I would generally be interested to hear your arguments/models in detail, if you get the chance to lay them out. My default position is LW has done well enough historically (e.g. Ben Pace's examples) for me to currently be investing in getting it even better. Epistemics and progress could definitely be a lot better, but getting there is hard. If I didn't see much progress on the rate of progress in the next year or two, I'd probably go focus on other things, though I think it'd be tragic if we ever lost what we have now. And another thought: Yes and no. Journal articles have their advantages, and so do blog posts. A bunch of recent LessWrong team's work has been around filling in the missing pieces for the system to work, e.g. Open Questions (hasn't yet worked for coordinating research), Annual Review, Tagging, Wiki. We often talk about conferences and "campus". My work on Open Questions involved thinking about i) a better template for articles than "Abstract, Intro, Methods, etc.", but Open Questions didn't work for unrelated reasons we haven't overcome yet, ii) getting lit reviews done systematically by people, iii) coordinating groups around research agendas. I've thought about re-attempting the goals of Open Questions with instead a "Research Agenda" feature that lets people communally maintain research agendas and work on them. It's a question of priorities whether I work on that anytime soon. I do really think many of the deficiencies of LessWrong's current work compared to academia are "infrastructure problems" at least as much as the epistemic standards of the community. Which means the LW team should be held culpable for not having solved them yet, but it is tricky.

7Richard_Ngo5y

For the record, I think the LW team is doing a great job. There's definitely a sense in which better infrastructure can reduce the need for high epistemic standards, but it feels like the thing I'm pointing at is more like "Many LW contributors not even realising how far away we are from being able to reliably produce and build on good ideas" (which feels like my criticism of Ben's position in his comment, so I'll respond more directly there).

5Pongo5y

It seems really valuable to have you sharing how you think we’re falling epistemically short and probably important for the site to integrate the insights behind that view. There are a bunch of ways I disagree with your claims about epistemic best practices, but it seems like it would be cool if I could pass your ITT more. I wish your attempt to communicate the problems you saw had worked out better. I hope there’s a way for you to help improve LW epistemics, but also get that it might be costly in time and energy.

4Viliam5y

Now they're positive again. Confusing to me, their Ω-karma (karma on another website) is also positive. Does it mean they previously had negative LW-karma but positive Ω-karma? Or that their Ω-karma also improved as a result of you complaining on LW a few hours ago? Why would it? (Feature request: graph of evolution of comment karma as a function of time.)

2Richard_Ngo5y

I'm confused, what is Ω-karma?

3MikkW5y

AI Alignment Forum karma (which is also displayed here on posts that are crossposted)

1[anonymous]5y

I'd be curious what, if any, communities you think set good examples in this regard. In particular, are there specific academic subfields or non-academic scenes that exemplify the virtues you'd like to see more of?

5Richard_Ngo5y

Maybe historians of the industrial revolution? Who grapple with really complex phenomena and large-scale patterns, like us, but unlike us use a lot of data, write a lot of thorough papers and books, and then have a lot of ongoing debate on those ideas. And then the "progress studies" crowd is an example of an online community inspired by that tradition (but still very nascent, so we'll see how it goes). More generally I'd say we could learn to be more rigorous by looking at any scientific discipline or econ or analytic philosophy. I don't think most LW posters are in a position to put in as much effort as full-time researchers, but certainly we can push a bit in that direction.

4[anonymous]5y

Thanks for your reply! I largely agree with drossbucket's reply. I also wonder how much this is an incentives problem. As you mentioned and in my experience, the fields you mentioned strongly incentivize an almost fanatical level of thoroughness that I suspect is very hard for individuals to maintain without outside incentives pushing them that way. At least personally, I definitely struggle and, frankly, mostly fail to live up to the sorts of standards you mention when writing blog posts in part because the incentive gradient feels like it pushes towards hitting the publish button. Given this, I wonder if there's a way to shift the incentives on the margin. One minor thing I've been thinking of trying for my personal writing is having a Knuth or Nintil style "pay for mistakes" policy. Do you have thoughts on other incentive structures to for rewarding rigor or punishing the lack thereof?

6Richard_Ngo5y

It feels partly like an incentives problem, but also I think a lot of people around here are altruistic and truth-seeking and just don't realise that there are much more effective ways to contribute to community epistemics than standard blog posts. I think that most LW discussion is at the level where "paying for mistakes" wouldn't be that helpful, since a lot of it is fuzzy. Probably the thing we need first are more reference posts that distill a range of discussion into key concepts, and place that in the wider intellectual context. Then we can get more empirical. (Although I feel pretty biased on this point, because my own style of learning about things is very top-down). I guess to encourage this, we could add a "reference" section for posts that aim to distill ongoing debates on LW. In some cases you can get a lot of "cheap" credit by taking other people's ideas and writing a definitive version of them aimed at more mainstream audiences. For ideas that are really worth spreading, that seems useful.

[-]Richard_Ngo2y*6810

Here is the best toy model I currently have for rational agents. Alas, it is super messy and hacky, but better than nothing. I'll call it the BAVM model; the one-sentence summary is "internal traders concurrently bet on beliefs, auction actions, vote on values, and merge minds". There's little novel here, I'm just throwing together a bunch of ideas from other people (especially Scott Garrabrant and Abram Demski).

In more detail, the three main components are:

A prediction market
An action auction
A value election

You also have some set of traders, who can simultaneously trade on any combination of these three. Traders earn money in two ways:

Making accurate predictions about future sensory experiences on the market.
Taking actions which lead to reward or increase the agent's expected future value.

They spend money in three ways:

Bidding to control the agent's actions for the next N timesteps.
Voting on what actions get reward and what states are assigned value.
Running the computations required to figure out all these trades.

Values are therefore dominated by whichever traders earn money from predictions or actions, who will disproportionately vote for values that are formulated in the same on... (read more)

8kave2y

I wonder if there's a loopiness here is which breaks the setup (the expectation I'm guessing is relative to the prediction markets probabilities? Though it seems like the market is over sensory experiences but the values are over world states in general, so maybe I'm missing something). But it seems like if I take an action and move the market at the same time, I might be able to extract a bunch of extra money and acquire outsize control. This seems like it's wasteful relative to contributing to a pool that bids on action A (or short-term policy P). I guess coordination is hard if you're just contributing to the pool though, and all connects to the merging process you describe.

4Nathan Helm-Burger2y

I've been studying and thinking about the physical side of this phenomenon in neuroscience recently. There are groups of columns of neurons in the cortex that form temporary voting blocks, regarding whatever subject that particular Brodmann area focuses on. These alternating groups have to deal with physical limits of how many groups the regions can stably divide into, which limits the number of active distinct hypotheses or 'traders' there can be in a given area at a given time. Unclear exactly what the max is, and it depends on the cortical region in question, but generally 6-9 is the approximate max (not coincidentally the number of distinct 'chunks' we can hold in active short term memory). Also, there is a tendency for noise to collapse too similar of traders/hypotheses/firing-groups to fall back into synchrony/agreement with each other and thus collapse back down to a baseline of two competing hypotheses. These hypotheses/firing-groups/traders are pushed into existence or pushed into merging not just by their own 'bids' but also by the evidence coming in from other brain areas or senses. I don't think that current day neuroscience has all the details yet (although I certainly don't have the full picture of all relevant papers in neuroscience!).

2Seth Herd2y

I think there are probably a lot of ways to build rational agents. The idea that general intelligence is hard in any absolute sense may be a biased by wanting to believe we're special, and for AI workers, that our work is special and difficult.

2Mateusz Bagiński2y

Can new traders be "spawned"?

4Richard_Ngo2y

Yepp, as in Logical Induction, new traders get spawned over time (in some kind of simplicity-weighted ordering).

2mako yass2y

I note that AI economies like this will often have explosively better credit assignment for information production than human economies can. Artificial agents can be copied or rolled back (erase memories), which makes it possible to reverse the receipt of information if an assessor concludes with a price that the seller considers too low for a deal. In human economies, that's impossible, you can't send a clone to value a piece of information then delete them if you decide not to buy that information (that's too expensive/illegal) nor can you wipe their memory of the information (or, we don't know how to do that), so the very basic requirement for trade, assessment prior to purchase, is not possible in human economies, so information doesn't get priced accurately and it has to be treated as a public good. When implementing this (internal privacy) in a multi-agent architecture, though, make sure to take measures to prevent the formation of monopolies, I feel like information is kind of an increasing returns type of good, yeah? The more you have the more you can do with it. It could quickly stop being multi-agent, and at worst, the monopoly could consolidate enough political power to manipulate the EV estimators and reward hack. In theory those economies shouldn't interact. But it's impossible to totally prevent it. The EV estimators are receiving big sets of action proposals from the decisionmakers and the decisionmakers will see which action proposal the EV estimators end up choosing.

2Richard_Ngo2y

Yepp, very good point. Am working on a short story about this right now.

2Daniel Kokotajlo2y

Nice! For comparison, this "Great Map of the Mind" is basically the standard academic philosophy picture.

2Vladimir_Nesov2y

My guess is that understanding merging is the key to most prediction-of-behavior issues (things that motivated and also foiled UDT, but not limited to known-in-advance preference setting). Two agents can coordinate if they are the same, or reasoning about each other's behavior, but in general they can be too complicated to clearly understand each other or themselves, can inadvertently diagonalize such attempts into impossibility, or even fail to be sufficiently aware of each other to start reasoning about each other specifically. It might be useful to formulate smaller computations (contracts/adjudicators) that facilitate coordination between different agents by being shared between them, with the bigger agents acting as parts of environments for the contracts and setting up incentives for them, while the contracts can themselves engage in decision making within those environments. Contracts coordinate by being shared and acting with strategicness across relevant agents (they should be something like common knowledge), and it's feasible for agents to find/construct some shared contracts as a result of them being much simpler than agents that host them. Learning of contracts doesn't need to start with targeting coordination with other big agents, as active contracts screen off the other agents they facilitate coordination with. Using contracts requires the big agents to make decisions about policies that affect the contracts updatelessly with respect to how the contracts end up behaving. That is, a contract should be able to know these policies, and the policies should describe responses to possible behaviors of a contract without themselves changing (once the contract computes more of its behavior), enabling the contract to do decision making in the environment of these policies. This corresponds to committing to abide by the contract. Assurance contracts (that start their tenure by checking that the commitments of all parties are actually in place) are especially

1Canaletto2y

If traders can get access to control panel for actions of the external agent AND they profit from accurately predicting its observations, then wouldn't the best strategy be "create as much chaos as possible that is only predictable to me, its creator". So, traders that value ONLY accurate predictions will get the advantage?

1Martín Soto2y

I like this picture! But I think real learning has some kind of ground-truth reward. So we should clearly separate between "this ground-truth reward that is chiseling the agent during training (and not after training)", and "the internal shards of the agent negotiating and changing your exact objective (which can happen both during and after training)". I'd call the latter "internal value allocation", or something like that. It doesn't neatly correspond to any ground truth, and is partly determined by internal noise in the agent. And indeed, eventually, when you "stop training" (or at least "get decoupled enough from reward"), it just evolves of its own, separate from any ground truth. And maybe more importantly: * I think this will by default lead to wireheading (a trader becomes wealthy and then sets reward to be very easy for it to get and then keeps getting it), and you'll need a modification of this framework which explains why that's not the case. * My intuition is a process of the form "eventually, traders (or some kind of specialized meta-traders) change the learning process itself to make it more efficient". For example, they notice that topic A and topic B are unrelated enough, so you can have the traders thinking about these topics be pretty much separate, and you don't lose much, and you waste less compute. Probably these dynamics will already be "in the limit" applied by your traders, but it will be the dominant dynamic so it should be directly represented by the formalism. * Finally, this might come later, and not yet in the level of abstraction you're using, but I do feel like real implementations of these mechanisms will need to have pretty different, way-more-local structure to be efficient at all. It's conceivable to say "this is the ideal mechanism, and real agents are just hacky approximations to it, so we should study the ideal mechanism first". But my intuition says, on the contrary, some of the physical constraints (like locality, or the

2Richard_Ngo2y

I'd actually represent this as "subsidizing" some traders. For example, humans have a social-status-detector which is hardwired to our reward systems. One way to implement this is just by taking a trader which is focused on social status and giving it a bunch of money. I think this is also realistic in the sense that our human hardcoded rewards can be seen as (fairly dumb) subagents. I think this happens in humans—e.g. we fall into cults, we then look for evidence that the cult is correct, etc etc. So I don't think this is actually a problem that should be ruled out—it's more a question of how you tweak the parameters to make this as unlikely as possible. (One reason it can't be ruled out: it's always possible for an agent to end up in a belief state where it expects that exploration will be very severely punished, which drives the probability of exploration arbitrarily low.) I'm assuming that traders can choose to ignore whichever inputs/topics they like, though. They don't need to make trades on everything if they don't want to. Yeah, this is why I'm interested in understanding how sub-markets can be aggregated into markets, sub-auctions into auctions, sub-elections into elections, etc.

1Martín Soto2y

Sounds good! Absolutely, wireheading is a real phenomenon, so the question is how can real agents exist that mostly don't fall to it. And I was asking for a story about how your model can be altered/expanded to make sense of that. My guess is it will have to do with strongly subsidizing some traders, and/or having a pretty weird prior over traders. Maybe even something like "dynamically changing the prior over traders"[1]. Yep, that's why I believe "in the limit your traders will already do this". I just think it will be a dominant dynamic of efficient agents in the real world, so it's better to represent it explicitly (as a more hierarchichal structure, etc.), instead of have that computation be scattered between all independent traders. I also think that's how real agents probably do it, computationally speaking. 1. ^ Of course, pedantically, yo will always be equivalent to having a static prior and changing your update rule. But some update rules are made sense of much easily if you interpret them as changing the prior.

2Richard_Ngo2y

Ah, I see. In that case I think I disagree that it happens "by default" in this model. A few dynamics which prevent it: 1. If the wealthy trader makes reward easier to get, then the price of actions will go up accordingly (because other traders will notice that they can get a lot of reward by winning actions). So in order for the wealthy trader to keep making money, they need to reward outcomes which only they can achieve, which seems a lot harder. 2. I don't yet know how traders would best aggregate votes into a reward function, but it should be something which has diminishing marginal return to spending, i.e. you can't just spend 100x as much to get 100x higher reward on your preferred outcome. (Maybe quadratic voting?) 3. Other traders will still make money by predicting sensory observations. Now, perhaps the wealthy trader could avoid this by making observations as predictable as possible (e.g. going into a dark room where nothing happens—kinda like depression, maybe?) But this outcome would be assigned very low reward by most other traders, so it only works once a single trader already has a large proportion of the wealth. IMO the best way to explicitly represent this is via a bias towards simpler traders, who will in general pay attention to fewer things. But actually I don't think that this is a "dominant dynamic" because in fact we have a strong tendency to try to pull different ideas and beliefs together into a small set of worldviews. And so even if you start off with simple traders who pay attention to fewer things, you'll end up with these big worldviews that have opinions on everything. (These are what I call frames here.)

1Martín Soto2y

Yep! But this didn't seem so hard for me to happen, especially in the form of "I pick some easy task (that I can do perfectly), and of course others will also be able to do it perfectly, but since I already have most of the money, if I just keep investing my money in doing it I will reign forever". You prevent this from happening through epsilon-exploration, or something equivalent like giving money randomly to other traders. These solutions feel bad, but I think they're the only real solutions. Although I also think stuff about meta-learning (traders explicitly learn about how they should learn, etc.) probably pragmatically helps make these failures less likely. Yep, that should help (also at the trade-off of making new good ideas slower to implement, but I'm happy to make that trade-off). Yeah. To be clear, the dynamic I think is "dominant" is "learning to learn better". Which I think is not equivalent to simplicity-weighing traders. It is instead equivalent to having some more hierarchichal structure on traders.

[-]Richard_Ngo1moΩ266418

Someone on the EA forum asked why I've updated away from public outreach as a valuable strategy. My response:

I used to not actually believe in heavy-tailed impact. On some gut level I thought that early rationalists (and to a lesser extent EAs) had "gotten lucky" in being way more right than academic consensus about AI progress. I also implicitly believed that e.g. Thiel and Musk and so on kept getting lucky, because I didn't want to picture a world in which they were actually just skillful enough to keep succeeding (due to various psychological blockers).

Now, thanks to dealing with a bunch of those blockers, I have internalized to a much greater extent that you can actually be good not just lucky. This means that I'm no longer interested in strategies that involve recruiting a whole bunch of people and hoping something good comes out of it. Instead I am trying to target outreach precisely to the very best people, without compromising much.

Relatedly, I've updated that the very best thinkers in this space are still disproportionately the people who were around very early. The people you need to soften/moderate your message to reach (or who need social proof in order to get involved)... (read more)

[-]Eli Tyre1mo224

The people you need to soften/moderate your message to reach (or who need social proof in order to get involved) are seldom going to be the ones who can think clearly about this stuff.

I strongly agree with this. (I wrote a post about it years ago.^[1])

Even of the people who were not "in early", of the ones who I most respect, and who seem to me to be doing the most impressive work that I'm most grateful to have in the world, 0 of them needed hand-holding or "outreach" to get them on board.

Writing the sequences was an amazing, high quality intervention that continues to pay dividends to this day. I think writing on the internet about the things that you think are important is a fantastic strategy, at least if your intellectual taste is good.

The payoff of most of the "movement building" and "community building" seems much murkier to me. At least some of it was clearly positive, but I don't know if it was positive on net (I think a smaller and more intense EA than the one we have in practice probably would have been better).

There's selection bias in kinds of community building I observed, but it seems to me that community building was more effective to the extent that it wa... (read more)

4Richard_Ngo1mo

Your post is great, I encourage you to repost it.

[-]plex1mo160

I agree with the main generator of this post (a small number of people produce a wildly disproportionate amount of the intellectual progress on hard problems) and one of the conclusions (don't water down your messages at all, if people need watered down messages they are unlikely to be helpful) but I think there's significant value in trying to communicate the hard problem of alignment broadly anyway because:

Filtering who are the best people is expensive and error-prone, so if you don't put the correct models in general circulation even pretty great people might just not become aware of them
People who are highly competent but not highly confident seem to often run into people who have been misinformed and become less sure of their own positions, having more generally circulating models of the main threat models would help those people get less distracted
Planting lots of seeds can be relatively cheap.

Also, related anecdote, I ran ~8 retreats at my house covering around 60 people in 2022/23. I got a decent read on how much of the core stack of alignment concepts at least half of them had, and how often they made hopeful mistakes which were transparently going to fail based on not hav... (read more)

[-]Cleo Nardo26d*Ω7142

Some thoughts on public outreach and "Were they early because they were good or lucky?"

Who are the best newcomers to AI safety? I'd be interested to here anyone's takes, not just Richard's. Who has done great work (by your lights) since joining after ChatGPT?
Rob Miles was the high watermark of public outreach. Unfortunately he stopped making videos. I'd be far more excited by a newcomer if they were persuaded by a Rob Miles video than an 80K video -- videos like 80K's "We're Not Ready for Superintelligence"^[1] are better on legible/easy-to-measure dimensions but worse in some more important way I think.
I observe a suspicious amount of 'social contagion' among the pre-ChatGPT AI Safety crowd, which updates me somewhat in favour of "lucky" over "good".^[2]

^{^}
^{^}
A bit anecdotal but: there are ~ a dozen people who went to our college in 2017-2020 now working full-time in AI safety, which is much higher than other colleges at the same university. I'm not saying any of us are particularly "great" -- but this suggests social contagion / information cascade, rather than "we figured this stuff out from the empty string". Maybe if you go back further (e.g. 2012-2016) there was less social co

... (read more)

[-]Richard_Ngo22dΩ184414

In trying to reply to this comment I identified four "waves" of AI safety, and lists of the central people in each wave. Since this is socially complicated I'll only share the full list of the first wave here, and please note that this is all based on fuzzy intuitions gained via gossip and other unreliable sources.

The first wave I’ll call the “founders”; I think of them as the people who set up the early institutions and memeplexes of AI safety before around 2015. My list:

Eliezer Yudkowsky
Michael Vassar
Anna Salamon
Carl Schulman
Scott Alexander
Holden Karnofsky
Nick Bostrom
Robin Hanson
Wei Dai
Shane Legg
Geoff Anders

The second wave I’ll call the “old guard”; those were the people who joined or supported the founders before around 2015. A few central examples include Paul Christiano, Chris Olah, Andrew Critch and Oliver Habryka.

Around 2014/2015 AI safety became significantly more professionalized and growth-oriented. Bostrom published Superintelligence, the Puerto Rico conference happened, OpenAI was founded, DeepMind started a safety team (though I don't recall exactly when), and EA started seriously pushing people towards AI safety. I’ll call the people who entered the field from then un... (read more)

[-]Amalthea21d101

My guess would be that nowadays many people who could bring a fresh perspective, or simply high-caliber original thinking, get either selected out/drowned out or are pushed through social and financial incentives to align there thinking towards more "mainstream" views.

8Wei Dai21d

Given that Vernor Vinge wrote The Coming Technological Singularity: How to Survive in the Post-Human Era in 1993, which single-handedly established much of the memeplex, including the still ongoing AI-first vs IA-first debate, another interesting question is why didn't anyone found the AI safety field until around 2000. For me, I'm not sure when I read this essay, but I did read Vinge's A Fire Upon the Deep in 1994 as a college freshman, which made me worried about a future AI takeover, but (as I wrote previously) I thought there would be plenty of smarter people working in AI safety so I went into applied cryptography instead (as a form of d/acc). Eliezer after reading Vinge (as a teen) didn't immediately heed the implicit or explicit safety warnings and instead wanted to accelerate the arrival of the Singularity as much as possible. It took him until around 2000 to pivot to safety. Nick Bostrom I think was concerned from the beginning or very early, but he was a PhD student when he got interested and I guess it took him a while to work through the academic system until he could found FHI in 2005. Maybe the real question is why didn't anyone else, i.e., someone with established credentials and social capital, found the field. Why did the task fall to a bunch of kids/students? The fact that nobody did it earlier does seem to suggest that it takes a very rare confluence of factors/circumstances for someone to do it. (Another tangential puzzle is why Vinge himself didn't get involved, as he was a professor of computer science in addition to science fiction writer. AFAIK he stayed completely off the early mailing lists as well as OB/LW nor had any contacts with anyone in AI safety.)

2Richard_Ngo15d

I'm not surprised by this, my sense is that it's usually young people and outsiders who pioneer new fields. Older people are just so much more shaped by existing paradigms, and also have so much more to lose, that it outweighs the benefits of their expertise and resources. Also 1993 to 2000 doesn't seem like that large a gap to me. Though I guess the thing I'm pointing at could also be summarized as "why hasn't someone created a new paradigm of AI safety in the last decade?" And one answer is that Paul and Chris and a few others created a half-paradigm of "ML safety", but it hasn't yet managed to show impressive enough results to fully take over. However, it did win on a memetic level amongst EAs in particular. The task at hand might then be understood as synthesizing the original "AI safety" with "ML safety". Or, to put it a bit more poetically, it's synthesizing the rationalist approach to aligning AGI with the empiricist approach to aligning AGI.

3Wei Dai15d

All of the fields that come to my mind (cryptography, theory of computation, algorithmic information theory, decision theory, game theory) were founded by much more established researchers. (But on reflection these all differ from AI safety by being fairly narrow and technical/mathematical, at least at their founding.) Which fields are you thinking of, that were founded by younger people and outsiders? Perplexity AI Pro (with GPT-5.1-Thinking)'s answer to "Who were the founders of academic cryptography research as a field and what where their jobs at the time?" There isn’t a single universally agreed-on “founder” of academic cryptography. Instead, a small group of researchers in the 1940s–1970s are usually credited with turning cryptography into an open, university-based research field. No single founder Histories of the subject generally describe a progression: Claude Shannon’s mathematical theory of secrecy in the 1940s, followed by the public‑key revolution of the 1970s and early 1980s that created today’s academic cryptography community. Shannon’s work was foundational, but it did not yet create an academic field in the modern sense; that came later with Whitfield Diffie, Martin Hellman, Ralph Merkle, and the inventors of RSA, whose work is often described as pioneering “modern” cryptography and has been recognized by ACM Turing Awards for cryptography pioneers.wikipedia+1 Early mathematical groundwork Claude Shannon is widely regarded as the founder of mathematical cryptography; in the 1940s he worked at Bell Labs as a researcher, where he developed the information‑theoretic framework for secrecy systems that later influenced public‑key cryptography. At roughly the same time and into the 1960s, cryptography research also existed in industry—most notably at IBM, where Horst Feistel headed an internal cryptography research group that designed ciphers such as Lucifer, which evolved into the Data Encryption Standard (DES), but this work was largely not yet

2Cole Wyeth15d

Chaitin was quite young when he (co-)invented AIT.

3Amalthea21d

Basically I'd bet capable people are still around, only that the circumstances don't allow them to rise to the top for whatever reason.

[-]Random Developer1mo*135

Being able to take future AI seriously as a risk seems to be highly correlated to being able to take COVID seriously as a risk in February 2020.

The key skill here may be as simple as being able to selectively turn off normalcy bias in the face of highly unusual news.
A closely related "skill" may be a certain general pessimism about future events, the sort of thing economists jokingly describe as "correctly predicting 12 out of the last 6 recessions."

That said, mass public action can be valuable. It's a notoriously blunt tool, though. As one person put it, "if you want to coordinate more than 5,000 people, your message can be about 5 words long." And the public will act anyway, in some direction. So if there's something you want to public to do, it can be worth organizing and working on communication strategies.

My political platform is, if you boil it down far enough, about 3 words long: "Don't build SkyNet." (As early as the late 90s, I joked about having a personal 11th commandment: "Thou shalt not build SkyNet." One of my career options at that point was to potentially work on early semi-autonomous robotic weapon platform prototypes, so this was actually relevant moral advic... (read more)

[-]dr_s1mo107

Yeah, it's not like the point of outreach is to mobilise citizen science on alignment (though that may happen). It's because in democracy the public is an important force. You can pick the option of focusing on converting a few powerful people and hope they can get shit done via non-political avenues but that hasn't worked spectacularly either for now, such people are still subject to classic race to the bottom dynamics and then you get cases like Altman and Musk, who all in all may have ended up net negative for the AI safety cause.

5Yair Halberstadt1mo

What about persuading politicians that AI safety is a cause that will win them votes? That requires very broad spectrum outreach to get as many ordinary people on board as possible.

4Wei Dai24d

From the context on EA Forum it seems clear that by "public outreach" you meant outreach to potential researchers to interest them in doing AI safety research, whereas a lot of people here seem to have misinterpreted your comment to have a broader meaning, to include, e.g., outreach to politicians and voters to try to influence future government policies.

3Oliver Daniels1mo

Part of the subtext here being the very best people (on the relevant dimensions) will naturally run into lesswrong, x-risk, etc, such that "out-reach" (in the sense of uni-organizing, advertising, etc) isn't that valuable on the current margin. To "the very best", doing high quality research is often the best "out-reach".

1OVERmind1mo

I think this is true only in a part of contexts. If we are talking about AI alignment - probably skilled mathematicians or AI researches can be very fit. At least in the directions like interpretability. And this doesn’t necessarily correlate with their desires to do societally unconventional work. Why isn’t that so?

1Bogdan Ionut Cirstea1mo

This seems to assume that the quality of labor of a small, highly-selected number of researchers, can be more important than a much larger amount of somewhat lower-quality labor, from a much larger number of participants. Seems like a pretty dubious assumption, especially given that other strategies seem possible. E.g. using a larger pool of participants to produce more easily verifiable, more prosaic AI safety research now, even at the risk of lower quality, so as to allow for better alignment + control of the kinds of AI models which will in the future for the first time be able to automate the higher quality and maybe less verifiable (e.g. conceptual) research that fewer people might be able to produce today. Put more briefly: quantity can have a quality of its own, especially in more verifiable research domains. Some of the claims around the quality of early rationalist / EA work also seem pretty dubious. E.g. a lot of the Yudkowsky-and-friends worldview is looking wildly overconfident and likely wrong.

[-]Richard_Ngo2y*Ω2662-13

Some opinions about AI and epistemology:

One reasons that many rationalists have such strong views about AI is that they are wrong about epistemology. Specifically, bayesian rationalism is a bad way to think about complex issues.
A better approach is meta-rationality. To summarize one guiding principle of (my version of) meta-rationality in a single sentence: if something doesn't make sense in the context of group rationality, it probably doesn't make sense in the context of individual rationality either.
For example: there's no privileged way to combine many people's opinions into a single credence. You can average them, but that loses a lot of information. Or you can get them to bet on a prediction market, but that depends on a lot on details of the individuals' betting strategies. The group might settle on a number to help with planning and communication, but it's only a lossy summary of many different beliefs and models. Similarly, we should think of individuals' credences as lossy summaries of different opinions from different underlying models that they have.
How does this apply to AI? Suppose we each think of ourselves as containing many different subagents that focus on u

... (read more)

[-]habryka2y*Ω123743

That will push P(doom) lower because most frames from most disciplines, and most styles of reasoning, don't predict doom.

I don't really buy this statement. Most frames, from most disciplines, and most styles of reasoning, do not make clear predictions about what will happen to humanity in the long-run future. A very few do, but the vast majority are silent on this issue. Silence is not anything like "50%".

Most frames, from most disciplines, and most styles of reasoning, don't predict sparks when you put metal in a microwave. This doesn't mean I don't know what happens when you put metal in a microwave. You need to at the very least limit yourself to applicable frames, and there are very few applicable frames for predicting humanity's long-term future.

4Joe Collman2y

I agree with this. Unfortunately, I think there's a fundamentally inside-view aspect of [problems very different from those we're used to]. I think looking for a range of frames is the right thing to do - but deciding on the relevance of the frame can only be done by looking at the details of the problem itself (if we instead use our usual heuristics for relevance-of-frame-x, we run into the same out-of-distribution issues). I don't think there's a way around this. Aspects of this situation are fundamentally different from those we're used to. [Is different from] is not a useful relation - we can't get far by saying "We've seen [fundamentally different] situations before - what happened there?". It'll all come back to how they were fundamentally different. To say something mildly more constructive, I do still think we should be considering and evaluating other frames, based on our own inside-view model (with appropriate error bars on that model). A place I'd start here would be: * Attempt to understand another frame. * See how far I need to zoom out before that frame's models become a reasonable abstraction for the problem-as-I-understand-it. * Find the smallest changes to my models that'd allow me to stick with this frame without zooming out so far. Assess the probability that these adjusted models are correct/useful. For most frames, I end up needing to zoom out too far for them to say much of relevance - so this doesn't much change my p(doom) assessment. It seems more useful to apply other frames to evaluate smaller parts of our models. I'm sure there are a bunch of places where intuitions and models from e.g. economics or physics do apply to safety-related subproblems.

[-]Wei Dai2y200

I've been thinking lately that human group rationality seems like such a mess. Like how can humanity navigate a once in a lightcone opportunity like the AI transition without doing something very suboptimal (i.e., losing most of potential value), when the vast majority of humans (and even the elites) can't understand (or can't be convinced to pay attention to) many important considerations. This big picture seems intuitively very bad and I don't know any theory of group rationality that says this is actually fine.
I guess my 1 is mostly about descriptive group rationality, and your 2 may be talking more about normative group rationality. However I'm also not aware of any good normative theories about group rationality. I started reading your meta-rationality sequence, but it ended after just two posts without going into details.
The only specific thing you mention here is "advance predictions" but for example, moral philosophy deals with "ought" questions and can't provide advance predictions. Can you say more about how you think group rationality should work, especially when advance predictions isn't possible?
From your group rationality perspective, why is it good that rationalists individually have better views about AI? Why shouldn't each person just say what they think from their own preferred frame, and then let humanity integrate that into some kind of aggregate view or outcome, using group rationality?

4mesaoptimizer2y

David Chapman's website seems like the standard reference for what the post-rationalists call "metarationality". (I haven't read much of it, but the little I read made me somewhat unenthusiastic about continuing).

[-]Thomas Kwa2yΩ3130

How can the mistakes rationalists are making be expressed in the language of Bayesian rationalism? Priors, evidence, and posteriors are fundamental to how probability works.

[-]Richard_Ngo1yΩ5106

The mistakes can (somewhat) be expressed in the language of Bayesian rationalism by doing two things:

Talking about partial hypotheses rather than full hypotheses. You can't have a prior over partial hypotheses, because several of them can be true at once (though you can still assign them credences and update those credences according to evidence).
Talking about models with degrees of truth rather than just hypotheses with degrees of likelihood. E.g. when using a binary conception of truth, general relativity is definitely false because it's inconsistent with quantum phenomena. Nevertheless, we want to say that it's very close to the truth. In general this is more of an ML approach to epistemology (we want a set of models with low combined loss on the ground truth).

[-]Max H2y110

Suppose we think of ourselves as having many different subagents that focus on understanding the world in different ways - e.g. studying different disciplines, using different styles of reasoning, etc. The subagent that thinks about AI from first principles might come to a very strong opinion. But this doesn't mean that the other subagents should fully defer to it (just as having one very confident expert in a room of humans shouldn't cause all the other humans to elect them as the dictator). E.g. maybe there's an economics subagent who will remain skeptical unless the AI arguments can be formulated in ways that are consistent with their knowledge of economics, or the AI subagent can provide evidence that is legible even to those other subagents (e.g. advance predictions).

Do "subagents" in this paragraph refer to different people, or different reasoning modes / perspectives within a single person? (I think it's the latter, since otherwise they would just be "agents" rather than subagents.)

Either way, I think this is a neat way of modeling disagreement and reasoning processes, but for me it leads to a different conclusion on the object-level question of AI doom.

A big part of why I f... (read more)

9Andrew_Critch2y

I may be missing context here, but as written / taken at face value, I strongly agree with the above comment from Richard. I often disagree with Richard about alignment and its role in the future of AI, but this comment is an extremely dense list of things I agree with regarding rationalist epistemic culture.

[-]mesaoptimizer2y105

I'd love to read an elaboration of your perspective on this, with concrete examples, which avoids focusing on the usual things you disagree about (pivotal acts vs. pivotal processes, social facets of the game is important for us to track, etc.) and mainly focus on your thoughts on epistemology and rationality and how it deviates from what you consider the LW norm.

8Noosphere891y

My main take on Bayesian epistemology being wrong is that I think to the extent it's useless in real life, it's because it focuses way too much on the ideal case, ala @Robert Miles's tweet here: https://x.com/robertskmiles/status/1830925270066286950 (The other problem I have with it is that even in the ideal case, it doesn't have a way to sensibly handle 0 probability events, or conditioning on probability 0 events, which can actually happen once we leave the world of finite sets and measures.) That said, I don't think that people being wrong about epistemology is the cause of high p(Doom). I'd agree more with @Algon in that the issues lie elsewhere (though a nitpick is that I wouldn't say that EU maximization is wrong for TAI/AGI/ASI, but rather that certain dangerous properties don't automatically hold, and that systems that EU maximize IRL like GPT-4 aren't actually nearly as dangerous as often assumed. Agree with the other points.)

8RobertM1y

(I am not the iniminatable @Robert Miles, though we do have some things in common.)

2Noosphere891y

Reply to @Algon: What I was talking about is that the predictive models like GPT-4 have a utility function that's essentially predictive, and the maximization is essentially trying to update the best it can given input conditions. These posts can help you to understand more about predictive/simulator utility functions like GPT-4: https://www.lesswrong.com/posts/vs49tuFuaMEd4iskA/one-path-to-coherence-conditionalization https://www.lesswrong.com/posts/k48vB92mjE9Z28C3s/implied-utilities-of-simulators-are-broad-dense-and-shallow https://www.lesswrong.com/posts/EBKJq2gkhvdMg5nTQ/instrumentality-makes-agents-agenty

2Algon1y

I'm doubtful that GPT-4 has a utility function. If it did, I would be kind-of terrified. I don't think I've seen the posts you linked to though, so I'll go read those.

2Noosphere891y

Maybe a crux is that I'm willing to grant learned utility functions as utility functions, and I tend to see EU maximization/utility function reasoning in general as implying far less consequences than people on LW think it is, at least without more constraints. It doesn't try to assert it's own existence, because that's not necessary for maximizing updating/prediction output based on inputs.

2Algon1y

I think the crux lies elsewhere, as I was sloppy in my wording. It's not that maximizing some utility function is an issue, as basically anything can be viewed as EU maximization for a sufficiently wild utility function. However, I don't view that as a meaningful utility function. Rather, it is the ones like e.g. utility functions over states that I think are meaningful, and those are scary. That's how I think you get classical paperclip maximizers. When I try and think up a meaningful utility function for GPT-4, I can't find anything that's plausible. Which means I don't think there's a meaningful prediction-utility function which describes GPT-4's behaviour. Perhaps that is a crux.

5Noosphere891y

Re utility functions over states, it turns out that we can validly turn utility functions over plans/predictions into utility functions over world states/outcomes (though usually with constraints on how large the domain is, though not always.) https://www.lesswrong.com/posts/k48vB92mjE9Z28C3s/?commentId=QciMJ9ehR9xbTexcc And yeah, I think it's a crux that I think that at the very least, what GPT-N systems will look like, if they reach AGI/ASI, will probably look like a maximizer for updating given input conditions like prompts. My main point isn't that the utility function framing of GPT-4 or GPT-N is wrong, but rather that LWers inferred way too much from how a system would behave, even conditional on expected utility maximization being a coherent frame for AIs, because they don't logically imply the properties they thought it did without more assumptions that need to be defended.

4MondSemmel2y

What is the empirical track record of your suggested epistemological strategy, relative to Bayesian rationalism? Where does your confidence come from that it would work any better? Every time I see suggestions of epistemological humility, I think to myself stuff like this: 1. What predictions would this strategy have made about future technologies, like an 1890 or 1900 prediction of the airplane (vs. first controlled flight by the Wright Brothers in 1903), or a 1930 or 1937 prediction of nuclear bombs? Doesn't your strategy just say that all these weird-sounding technologies don't exist yet and are probably impossible? 2. Can this epistemological strategy correctly predict that present-day huge complex machines like airplanes can exist? They consist of millions of parts and require contributions of thousands or tens of thousand of people. Each part has a chance of being defective, and each person has a chance of making a mistake. Without the benefit of knowing that airplanes do indeed exist, doesn't it sound overconfident to predict that parts have an error rate of <1 in a million, or that people have an error rate of <1 in a thousand? But then the math says that airplanes can't exist, or should immediately crash. 3. Or to rephrase point 2 to reply to this part: "That will push P(doom) lower because most frames from most disciplines, and most styles of reasoning, don't predict doom." — Can your epistemological strategy even correctly make any predictions of near 100% certainty? I concur with habryka that most frames don't make any predictions on most things. And yet this doesn't mean that some events aren't ~100% certain.

4quetzal_rainbow2y

One of the most important features of future ASI I consider knowledge of limits of applicability of its models and heuristics. If you have list of assumptions for very fast heuristics, then you can win big by doing fast-computable moves in narrow environment where assumptions hold. Thus saying, you need to be able find when your assumptions don't hold and command your subagents to halt, melt and catch fire when they are outside of their applicability zone.

3Algon2y

I think this post doesn't really explain why rats have high belief in doom, or why they're wrong to do so. Perhaps ironically, there is a better a version of this post on both counts which isn't so focused on how rats get epistemology wrong and the social/meta-level consequences. A post which focuses on the object-level implications for AI of a theory of rationality which looks very different from the AIXI-flavoured rat-orthodox view. I say this because those sorts of considerations convinced me that we're much less likely to be buggered. I.e. I no longer believe EU maximization is/will be a good description by default of TAI or widely economically productive AGI, mildly superhuman AGI or even ASI, depending on the details. Which is partly due to a recognition that the arguments for EU maximization are weaker than I thought, arguments for LDT being convergent are lacking, the notions of optimality we do have are very weak, the existence and behaviour of GPT-4, Claude Opus etc. 6 seems too general a claim to me. Why wouldn't it work for 1% vs 10%, and likewise 0.1% vs 1% i.e. why doesn't this suggest that you should round down P(doom) to zero. Also, I don't even know what you mean by "most" here. Like, are we quantifying over methods of reasoning used by current AI researchers right now? Over all time? Over all AI researchers and engineers? Over everyone in the West? Over everyone who's ever lived? Etc. And it seems to me like you're implicitly privileging ways of combining these opinions that get you 10% instead of 1% or 90%, which is begging the question. Of course, you could reply that a P(doom) of 10% is confused, that isn't really your state of knowledge, lumping in all your sub-agents models into a single number is too lossy etc. But then why mention that 90% is a much stronger prediction than 10% instead of saying they're roughly equally confused? 7 I kinda disagree with. Those models of idealized reasoning you mention generalize Bayesianism/Expected Ut

4Richard_Ngo2y

Thanks for the reply. I'm working on this right now, actually. Will hopefully post in a couple of weeks. That seems reasonable. But I do think there's a group of people who have internalized bayesian rationalism enough that the main blocker is their general epistemology, rather than the way they reason about AI in particular. I think the point of 6 is not to say "here's where you should end up", but more to say "here's the reason why this straightforward symmetry argument doesn't hold". There's still something importantly true about EU maximization and bayesianism. I think the changes we need will be subtle but have far-reaching ramifications. Analogously, relativity was a subtle change to newtonian mechanics that had far-reaching implications for how to think about reality. Any epistemology will rule out some updates, but a problem with bayesianism is that it says there's one correct update to make. Whereas radical probabilism, for example, still sets some constraints, just far fewer.

2Algon2y

This sounds cool. I think your OP didn't give enough details as to why internalizing Bayesian rationalism leads to doominess by default. Like, Nora Belrose is firmly Bayesian and is decidedly an optimist. Admittedly, I think she doesn't think a Kolmogorov prior is a good one, but I don't think that makes you much more doomy either. I think Jacob Cannel and others are also Bayesian and non-doomy. Perhaps I'm using "Bayesian rationalism" differently than you are, which is why I think your claim, as I read it, is invalid. Fair enough. However, how big is the asymmetry? I'm a bit sceptical there is a large one. Based off my interactions, it seems like ~ everyone who has seriously thought about this topic for a couple of hours has radically different models, w/ radically different levels of doominess. This holds even amongst people who share many lenses (e.g. Tyler Cowen vs Robin Hanson, Paul Christiano vs. Scott Aaronson, Steve Hsu vs Michael Nielsen etc.). I think we're in agreement over this. (I think Bayesianism less wrong than EU maximization, and probably a very good approximation in lots of places, like Newtonian physics is for GR.) But my contention is over Bayesian epistemology tripping many rats up when thinking about AI x-risk. You need some story which explains why sticking to Bayesian epistemology is tripping up very many people here in particular. Right, but in radical probabilism the type of beliefs is still a real valued function, no? Which is in tension w/ many disparate models that don't get compressed down to a single number. In that sense, the refined formalism is still rigid in a way that your description is flexible. And I suspect the same is true for Infra-Bayesianism, though I understand that even less well than radical probabilism.

2Seth Herd2y

I think you're making a good point (rationalists maybe don't weight other opinions highly enough), but you'd get farther framing it as an update to how to use Bayesian reasoning, rather than an alternative. Bayesian reasoning has a pretty strong intuitive connection to "the factually correct way to reason", even though there's a ton of subtlety in that statement and how and where it's applied. WRT to many of your arguments: base rates are increasingly just the wrong way to reason about AGI risks. We can think in more detail about how we'll build AGI and what the risks are.

1tmeanen2y

Am I misunderstandng this sentence? How do "90% doom" and the assumption that survival is the default square with one another?

2Richard_Ngo2y

Edited for clarity now.

1CstineSublime2y

I think they are just using that as an example of a strongly opinionated sub-agent which may be one of many different and highly specific probability assessments of doom. As for "survival is the default assumption" - what a declaration of that implies on the surface level is that the chance of survival is overwhelming except in the case of a cataclysmic AI scenario. To put it another way: we have a 99% chance of survival so long as we get AGI right. To put it yet another way - Hollywood has made popular films about the human world being destroyed by Nuclear War, Climate Change, Viral Pandemic, and Asteroid Impact to name a few - different sub-agents could each give higher or lower probabilities to each of those scenarios depending on things like domain knowledge and in concert it raises the question of why we presume that survival is the default? What is the ensemble average of doom? Is doom more or less likely than survival for any given time frame?

2[comment deleted]2y

[-]Richard_Ngo1mo*578

In a post on Solomonoff Induction (and also in this wiki entry), Yudkowsky describes Shannon’s minimax algorithm for searching the entire chess game-tree as an example of conceptual progress. Previously, Edgar Allen Poe had argued that it was impossible in principle for a machine to play chess well. With Shannon’s algorithm, it became possible in principle, just computationally infeasible.

However, even “principled” algorithms like minimax search don’t take into account the possibility that your opponent knows things you don’t know (as I’ve previously discussed here). And so there are many elements of good chess strategy for bounded agents that they can’t account for:

When your opponent plays a move that seems surprisingly bad, update that they’re probably seeing some tactics that you’re missing.
When your opponent plays a move that seems surprisingly good, update that they’re probably more skilled than you thought.
Try to play tricky (“sharp”) lines when you’re behind, in the hope that your opponent makes a mistake.
Try to play solidly when you’re ahead, even if it narrows your lead.
Identify your opponent’s playing style, and try to exploit it.

Some of these considerations have been dis... (read more)

6Mateusz Bagiński1mo

It seems to me like you're trying to solve a different problem. Unbounded minimax should handle all of this (in the sense that it won't be an obstacle). Unless you are talking about bounded approximations.

4Richard_Ngo1mo

Yeah, "incomplete" is the wrong word here. I edited shortly after posting to say instead "there are many elements of good chess strategy for bounded agents that they can’t account for".

4Adrià Garriga-alonso1mo

Pedantically, it's grounded in the likelihood of LeelaZero wining from that position when playing itself. Stockfish NNUE is trained on LZ data. Previously they were completely human-crafted, grounded on repeated empirical testing of Stockfish and picking what works better. I suppose they're still grounded in the latter.

4morelight1mo

You seem to be locating two things here: 1. the Knightian uncertainty of local positions within a game (“given this exact configuration, how much does my outcome depend on unknown opponent type, unknown payoff structure, unknown model of the world?”) and 2. the Knightian uncertainty of some games relative to other games ("how much does the rule set, state space, opponent ontology, or payoff structure mutate as play continues?”) I suspect that the two are related, and that when we calculate the Knightian uncertainty of local positions, we will need to factor in our uncertainty with regards to the type of game we're playing. For example: Tic-tac-toe has zero: the state space is fully enumerated, optimal play is strivial, every competent player forces a draw - there's no tactical complexity to amplify mistakes. Chess has slightly more meta-game permeability. Players cannot change the board, cannot add new pieces, cannot redefine legal moves, cannot declare “this now counts as checkmate”, cannot invent a new win condition. Go has the same fixed-rule structure, but the state space explodes. The game-tree is too large to fully reason about, so bounded reasoning dominates. In poker/markets, rules are mostly fixed, but information asymmetry is now part of the actual game. In diplomacy/geopolitics, rules are soft, not hard. Players can redefine objectives and also alter the payoff structure. The game itself can change while you are playing it. Art has even more meta-game permeability (and therefore a higher game_uncertainty coefficient): artists invent new art forms, collapse genres, redefine what counts as art, etc. Every innovation becomes part of the game for everyone else. (Publishers in 1914-1939 to James Joyce: "This isn't a novel." James Joyce: "Fuck you yes it is") Romance is the most Knightian domain maybe humans ever engage in. Signals are ambiguous, no stable payoff function, opponents mututate as agents (psychologically and physically)

1bodry1mo

I'll take a shot. If EAlice(mvs,Bob) is the expected return (1 for win, 0.5 for draw and 0 for loss) for Alice given that she's playing Bob, she knows Bob's source code, and the moves mvs have been played so far, then the value of a position for Alice is ∑pϵSEAlice(mvs,p)∗P(p|mvs) where S is the set of programs that Alice's opponent is drawn from.

[-]Richard_Ngo3yΩ26480

(Written quickly and not very carefully.)

I think it's worth stating publicly that I have a significant disagreement with a number of recent presentations of AI risk, in particular Ajeya's "Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover", and Cohen et al.'s "Advanced artificial agents intervene in the provision of reward". They focus on policies learning the goal of getting high reward. But I have two problems with this:

I expect "reward" to be a hard goal to learn, because it's a pretty abstract concept and not closely related to the direct observations that policies are going to receive. If you keep training policies, maybe they'd converge to it eventually, but my guess is that this would take long enough that we'd already have superhuman AIs which would either have killed us or solved alignment for us (or at least started using gradient hacking strategies which undermine the "convergence" argument). Analogously, humans don't care very much at all about the specific connections between our reward centers and the rest of our brains - insofar as we do want to influence them it's because we care about much more directly-observable p

... (read more)

[-]paulfchristiano3yΩ19292

I'm not very convinced by this comment as an objection to "50% AI grabs power to get reward." (I find it more plausible as an objection to "AI will definitely grab power to get reward.")

I expect "reward" to be a hard goal to learn, because it's a pretty abstract concept and not closely related to the direct observations that policies are going to receive
"Reward" is not a very natural concept

This seems to be most of your position but I'm skeptical (and it's kind of just asserted without argument):

The data used in training is literally the only thing that AI systems observe, and prima facie reward just seems like another kind of data that plays a similarly central role. Maybe your "unnaturalness" abstraction can make finer-grained distinctions than that, but I don't think I buy it.
If people train their AI with RLDT then the AI is literally be trained to predict reward! I don't see how this is remote, and I'm not clear if your position is that e.g. the value function will be bad at predicting reward because it is an "unnatural" target for supervised learning.
I don't understand the analogy with humans. It sounds like you are saying "an AI system selected based on the reward of its acti

... (read more)

2TurnTrout3y

I don't know what this means. Suppose we have an AI which "cares about reward" (as you think of it in this situation). The "episode" consists of the AI copying its network & activations to another off-site server, and then the original lab blows up. The original reward register no longer exists (it got blown up), and the agent is not presently being trained by an RL alg. What is the "reward" for this situation? What would have happened if we "sampled" this episode during training?

6paulfchristiano3y

I agree there are all kinds of situations where the generalization of "reward" is ambiguous and lots of different things could happen . But it has a clear interpretation for the typical deployment episode since we can take counterfactuals over the randomization used to select training data. It's possible that agents may specifically want to navigate towards situations where RL training is not happening and the notion of reward becomes ambiguous, and indeed this is quite explicitly discussed in the document Richard is replying to. As far as I can tell the fact that there exist cases where different generalizations of reward behave differently does not undermine the point at all.

2TurnTrout3y

Yeah, I think I was wondering about the intended scoping of your statement. I perceive myself to agree with you that there are situations (like LLM training to get an alignment research assistant) where "what if we had sampled during training?" is well-defined and fine. I was wondering if you viewed this as a general question we could ask. I also agree that Ajeya's post addresses this "ambiguity" question, which is nice!

2Richard_Ngo3y

It's intended as an objection to "AI grabs power to get reward is the central threat model to focus on", but I think our disagreements still apply given this. (FWIW my central threat model is that policies care about reward to some extent, but that the goals which actually motivate them to do power-seeking things are more object-level.) I expect policies to be getting rich input streams like video, text, etc, which they use to make decisions. Reward is different from other types of data because reward isn't actually observed as part of these input streams by policies during episodes. This makes it harder to learn as a goal compared with things that are more directly observable (in a similar way to how "care about children" is an easier goal to learn than "care about genes"). I don't think this line of reasoning works, because "the episode appearing in training" can be a dependent variable. For example, consider an RL agent that's credibly told that its data is not going to be used for training unless it misbehaves badly. An agent which maximizes reward conditional on the episode appearing in training is therefore going to misbehave the minimum amount required to get its episode into the training data (and more generally, behave like getting its episode into training is a terminal goal). This seems very counterintuitive. Some versions that wouldn't result in power-grabbing: * Goal is "get highest proportion of possible reward"; the policy might rewrite the training algorithm to be myopic, then get perfect reward for one step, then stop. * Goal is "care about (not getting low rewards on) specific computers used during training"; the policy might destroy those particular computers, then stop. * Goal is "impress the critic"; the policy might then rewrite its critic to always output high reward, then stop. * Goal is "get high reward myself this episode"; the policy might try to do power-seeking things but never create more copies of itself, and eventually lose co

5paulfchristiano3y

Are you imagining that such systems get meaningfully low reward on the training distribution because they are pursuing those goals, or that these goals are extremely well-correlated with reward on the training distribution and only come apart at test time? Is the model deceptively aligned? Children vs genes doesn't seem like a good comparison, it seems obvious that models will understand the idea of reward during training (whereas humans don't understand genes during evolution). A better comparison might be "have children during my life" vs "have a legacy after I'm gone," but in fact humans have both goals even though one is never directly observed. I guess more importantly, I don't buy the claim about "things you only get selected on are less natural as goals than things you observe during episodes," especially if your policy is trained to make good predictions of reward. I don't know if there's a specific reason for your view, or this is just a clash of intuitions. It feels to me like my position is kind of the default, in that you are offering a feature and saying it is a major consideration that SGD wouldn't learn a particular kind of cognition. If the agent misbehaves so that its data will be used for training, then the misbehaving actions will get a low reward. So SGD will shift from "misbehave so that my data will be used for training" to "ignore the effect of my actions on whether my data will be used for training, and just produce actions that would result in a low reward assuming that this data is used for training." I agree the behavior of the model isn't easily summarized by an English sentence. The thing that seems most clear is that the model trained by SGD will learn not to sacrifice reward in order to increase the probability that the episode is used in training. If you think that's wrong I'm happy to disagree about it. Every one of your examples results in the model grabbing power and doing something bad at deployment time, though I agree that

4Richard_Ngo3y

Reading back over this now, I think we're arguing at cross purposes in some ways. I should have clarified earlier that my specific argument was against policies learning a terminal goal of reward that generalizes to long-term power-seeking. I do expect deceptive alignment after policies learn other broadly-scoped terminal goals and realize that reward-maximization is a good instrumental strategy. So all my arguments about the naturalness of reward-maximization as a goal are focused on the question of which type of terminal goal policies with dangerous levels of capabilities learn first.* Let's distinguish three types (where "myopic" is intended to mean something like "only cares about the current episode"). 1. Non-myopic misaligned goals that lead to instrumental reward maximization (deceptive alignment) 2. Myopic terminal reward maximization 3. Non-myopic terminal reward maximization Either 1 and 2 (or both of them) seem plausible to me. 3 is the one I'm skeptical about. How come? 1. We should expect models to have fairly robust terminal goals (since, unlike beliefs or instrumental goals, terminal goals shouldn't change quickly with new information). So once they understand the concept of reward maximization, it'll be easier for them to adopt it as an instrumental strategy than a terminal goal. (An analogy to evolution: once humans construct highly novel strategies for maximizing genetic fitness (like making thousands of clones) people are more likely to do it for instrumental reasons than terminal reasons.) 2. Even if they adopt reward maximization as a terminal goal, they're more likely to adopt a myopic version of it than a non-myopic version, since (I claim) the concept of reward maximization doesn't generalize very naturally to larger scales. Above, you point out that even relatively myopic reward maximization will lead to limited takeover, and so we'll train subsequent agents to be less myopic. But it seems to me that the selection pressure generated

1Tom Davidson3y

Why does it not lead to takeover in the same way?

2paulfchristiano3y

Because it's easy to detect and correct (except that correcting it might push you into one of the other regimes).

1Tom Davidson3y

So far causally upstream of the human evaluator's opinion? Eg an AI counselor optimizing for getting to know you

0TurnTrout3y

(Emphasis added) I don't think this engages with the substance of the analogy to humans. I don't think any party in this conversation believes that human learning is "just" RL based on a reward circuit, and I don't believe it either. "Just RL" also isn't necessary for the human case to give evidence about the AI case. Therefore, your summary seems to me like a strawman of the argument. I would say "human value formation mostly occurs via RL & algorithms meta-learned thereby, but in the important context of SSL / predictive processing, and influenced by inductive biases from high-level connectome topology and genetically specified reflexes and environmental regularities and..." Furthermore, we have good evidence that RL plays an important role in human learning. For example, from The shard theory of human values:

4paulfchristiano3y

This is incredibly weak evidence. * Animals were selected over millions of generations to effectively pursue external goals. So yes, they have external goals. * Humans also engage in within-lifetime learning, so of course you see all kinds of indicators of that in brains. Both of those observations have high probability, so they aren't significant Bayesian evidence for "RL tends to produce external goals by default." In particular, for this to be evidence for Richard's claim, you need to say: "If RL tended to produce systems that care about reward, then RL would be significantly less likely to play a role in human cognition." There's some update there but it's just not big. It's easy to build brains that use RL as part of a more complicated system and end up with lots of goals other than reward. My view is probably the other way---humans care about reward more than I would guess from the actual amount of RL they can do over the course of their life (my guess is that other systems play a significant role in our conscious attitude towards pleasure).

2interstice3y

Curious what systems you have in mind here.

0TurnTrout3y

I don't understand why you think this explains away the evidential impact, and I guess I put way less weight on selection reasoning than you do. My reasoning here goes: 1. Lots of animals do reinforcement learning. 2. In particular, humans prominently do reinforcement learning. 3. Humans care about lots of things in reality, not just certain kinds of cognitive-update-signals. 4. "RL -> high chance of caring about reality" predicts this observation more strongly than "RL -> low chance of caring about reality" This seems pretty straightforward to me, but I bet there are also pieces of your perspective I'm just not seeing. But in particular, it doesn't seem relevant to consider selection pressures from evolution, except insofar as we're postulating additional mechanisms which evolution found which explain away some of the reality-caring? That would weaken (but not eliminate) the update towards "RL -> high chance of caring about reality." I don't see how this point is relevant. Are you saying that within-lifetime learning is unsurprising, so we can't make further updates by reasoning about how people do it? I'm saying that there was a missed update towards that conclusion, so it doesn't matter if we already knew that humans do within-lifetime learning?

9paulfchristiano3y

You seem to be saying P(humans care about the real world | RL agents usually care about reward) is low. I'm objecting, and claiming that in fact P(humans care about the real world | RL agents usually care about reward) is fairly high, because humans are selected to care about the real world and evolution can be picky about what kind of RL it does, and it can (and does) throw tons of other stuff in there. The Bayesian update is P(humans care about the real world | RL agents usually care about reward) / P(humans care about the real world | RL agents mostly care about other stuff). So if e.g. P(humans care about the real world | RL agents don't usually care about reward) was 80%, then your update could be at most 1.25. In fact I think it's even smaller than that.. And then if you try to turn that into evidence about "reward is a very hard concept to learn," or a prediction about how neural nets trained with RL will behave, it's moving my odds ratios by less than 10% (since we are using "RL" quite loosely in this discussion, and there are lots of other differences and complications at play, all of which shrink the update). You seem to be saying "yes but it's evidence," which I'm not objecting to---I'm just saying it's an extremely small amount evidence. I'm not clear on whether you agree with my calculation. (Some of the other text I wrote was about a different argument you might be making: that P(humans use RL | RL agents usually care about reward) is significantly lower than P(humans use RL| RL agents mostly are about other stuff), because evolution would then have never used RL. My sense is that you aren't making this argument so you should ignore all of that, sorry to be confusing.)

6TurnTrout3y

Just saw this reply recently. Thanks for leaving it, I found it stimulating. (I wrote the following rather quickly, in an attempt to write anything at all, as I find it not that pleasant to write LW comments -- no offense to you in particular. Apologies if it's confusing or unclear.) Yes, in large part. Yeah, are people differentially selected for caring about the real world? At the risk of seeming facile, this feels non-obvious. My gut take is that conditional on RL agents usually caring about reward (and thus setting aside a bunch of my inside-view reasoning about how RL dynamics work), conditional on that -- reward-humans could totally have been selected for. This would drive up P(humans care about reward | RL agents care about reward, humans were selected by evolution), and thus (I think?) drive down P(humans care about the real world | RL agents usually care about reward). POV: I'm in an ancestral environment, and I (somehow) only care about the rewarding feeling of eating bread. I only care about the nice feeling which comes from having sex, or watching the birth of my son, or being gaining power in the tribe. I don't care about the real-world status of my actual son, although I might have strictly instrumental heuristics about e.g. how to keep him safe and well-fed in certain situations, as cognitive shortcuts for getting reward (but not as terminal values). In what way is my fitness lower than someone who really cares about these things, given that the best way to get rewards may well be to actually do the things? Here are some ways I can think of: 1. Caring about reward directly makes reward hacking a problem evolution has to solve, and if it doesn't solve it properly, the person ends up masturbating and not taking (re)productive actions. 1. Counter-counterpoint: But also many people do in fact enjoy masturbating, even though it seems (to my naive view) like an obvious thing to select away, which was present ancestrally. 2. People seem t

7Vivek Hebbar3y

Would such a person sacrifice themselves for their children (in situations where doing so would be a fitness advantage)?

8TurnTrout3y

I think this highlights a good counterpoint. I think this alternate theory predicts "probably not", although I can contrive hypotheses for why people would sacrifice themselves (because they have learned that high-status -> reward; and it's high-status to sacrifice yourself for your kid). Or because keeping your kid safe -> high reward as another learned drive. Overall this feels like contortion but I think it's possible. Maybe overall this is a... 1-bit update against the "not selection for caring about reality" point?

[-]Richard_Ngo3yΩ14210

Putting my money where my mouth is: I just uploaded a (significantly revised) version of my Alignment Problem position paper, where I attempt to describe the AGI alignment problem as rigorously as possible. The current version only has "policy learns to care about reward directly" as a footnote; I can imagine updating it based on the outcome of this discussion though.

3dsj3y

For someone who's read v1 of this paper, what would you recommend as the best way to "update" to v3? Is an entire reread the best approach? [Edit March 11, 2023: Having now read the new version in full, my recommendation to anyone else with the same question is a full reread.]

1[comment deleted]3y

7Ajeya Cotra3y

Note that the "without countermeasures" post consistently discusses both possibilities (the model cares about reward or the model cares about something else that's consistent with it getting very high reward on the training dataset). E.g. see this paragraph from the above-the-fold intro: As well as the section Even if Alex isn't "motivated" to maximize reward.... I do place a ton of emphasis on the fact that Alex enacts a policy which has the empirical effect of maximizing reward, but that's distinct from being confident in the motivations that give rise to that policy. I believe Alex would try very hard to maximize reward in most cases, but this could be for either terminal or instrumental reasons. With that said, for roughly the reasons Paul says above, I think I probably do have a disagreement with Richard -- I think that caring about some version of reward is pretty plausible (~50% or so). It seems pretty natural and easy to grasp to me, and because I think there will likely be continuous online training the argument that there's no notion of reward on the deployment distribution doesn't feel compelling to me.

5Richard_Ngo3y

Yepp, agreed, the thing I'm objecting to is how you mainly focus on the reward case, and then say "but the same dynamics apply in other cases too..." The problem is that you need to reason about generalization to novel situations somehow, and in practice that ends up being by reasoning about the underlying motivations (whether implicitly or explicitly).

3Lauro Langosco3y

I agree with your general point here, but I think Ajeya's post actually gets this right, eg and

2Lauro Langosco3y

I also think that often "the AI just maximizes reward" is a useful simplifying assumption. That is, we can make an argument of the form "even if the AI just maximizes reward, it still takes over; if it maximizes some correlate of the reward instead, then we have even less control over what it does and so are even more doomed". (Though of course it's important to spell the argument out)

3Ajeya Cotra3y

Yeah, I agree this is a good argument structure -- in my mind, maximizing reward is both a plausible case (which Richard might disagree with) and the best case (conditional on it being strategic at all and not a bag of heuristics), so it's quite useful to establish that it's doomed; that's the kind of structure I was going for in the post.

9Richard_Ngo3y

I strongly disagree with the "best case" thing. Like, policies could just learn human values! It's not that implausible. If I had to try point to the crux here, it might be "how much selection pressure is needed to make policies learn goals that are abstractly related to their training data, as opposed to goals that are fairly concretely related to their training data?" Where we both agree that there's some selection pressure towards reward-like goals, and it seems like you expect this to be enough to lead policies to behavior that violates all their existing heuristics, whereas I'm more focused on the regime where there are lots of low-hanging fruit in terms of changes that would make a policy more successful, and so the question of how easy that goal is to learn from its training data is pretty important. (As usual, there's the human analogy: our goals are very strongly biased towards things we have direct observational access to!) Even setting aside this disagreement, though, I don't like the argumentative structure because the generalization of "reward" to large scales is much less intuitive than the generalization of other concepts (like "make money") to large scales - in part because directly having a goal of reward is a kinda counterintuitive self-referential thing.

4Ajeya Cotra3y

Yes, sorry, "best case" was oversimplified. What I meant is that generalizing to want reward is in some sense the model generalizing "correctly;" we could get lucky and have it generalize "incorrectly" in an important sense in a way that happens to be beneficial to us. I discuss this a bit more here. I don't understand why reward isn't something the model has direct access to -- it seems like it basically does? If I had to say which of us were focusing on abstract vs concrete goals, I'd have said I was thinking about concrete goals and you were thinking about abstract ones, so I think we have some disagreement of intuition here. Yeah, I don't really agree with this; I think I could pretty easily imagine being an AI system asking the question "How much reward would this episode get if it were sampled for training?" It seems like the intuition this is weird and unnatural is doing a lot of work in your argument, and I don't really share it.

5cfoster03y

AFAIK the reward signal is not typically included as an input to the policy network in RL. Not sure why, and I could be wrong about that, but that is not my main question. The bigger question is "Has direct access to when?" At the moment in time when the model is making a decision, it does not have direct access to the decision-relevant reward signal because that reward is typically causally downstream of the model's decision. That reward may not even have a definite value until after decision time. Whereas concrete observables like "shiny gold coins" and "the finish line straight ahead" and "my opponent is in check" (and other abstractions in the model's ontology that are causally upstream from reward in reality) are readily available at decision time. It seems to me that that makes them natural candidates for credit assignment to flag early on as the reward-responsible mental events and reinforce into stable motivations, since they in fact were the factors that determined the decisions that led to rewards. IME, the most straightforward way for reward-itself to become the model's primary goal would be if the model learns to base its decisions on an accurate reward-predictor much earlier than it learns to base its decisions on other (likely upstream) factors. If it instead learns how to accurately predict reward-itself after it is already strongly motivated by some concrete observables, I don't see why we should expect it to dislodge that motivation, despite the true fact that those concrete observables are only pretty correlated with reward whereas an accurate reward-predictor is perfectly correlated with reward. Why? Because the model currently doesn't care about reward-itself, it currently cares about the concrete observable(s), so it has no reason to take actions that would override that goal, and it has positive goal-content integrity reasons to not take those actions.

4TurnTrout3y

See also: Inner and outer alignment decompose one hard problem into two extremely hard problems (in particular: Inner alignment seems anti-natural).

[-]Richard_Ngo9moΩ234320

When you think of goals as reward/utility functions, the distinction between positive and negative motivations (e.g. as laid out in this sequence) isn’t very meaningful, since it all depends on how you normalize them.

But when you think of goals as world-models (as in predictive processing/active inference) then it’s a very sharp distinction: your world-model-goals can either be of things you should move towards, or things you should move away from.

This updates me towards thinking that the positive/negative motivation distinction is more meaningful than I thought.

[-]Steven Byrnes9mo334

In run-and-tumble motion, “things are going well” implies “keep going”, whereas “things are going badly” implies “choose a new direction at random”. Very different! And I suggest in §1.3 here that there’s an unbroken line of descent from the run-and-tumble signal in our worm-like common ancestor with C. elegans, to the “valence” signal that makes things seem good or bad in our human minds. (Suggestively, both run-and-tumble in C. elegans, and the human valence, are dopamine signals!)

So if some idea pops into your head, “maybe I’ll stand up”, and it seems appealing, then you immediately stand up (the human “run”); if it seems unappealing on net, then that thought goes away and you start thinking about something else instead, semi-randomly (the human “tumble”).

So positive and negative are deeply different. Of course, we should still call this an RL algorithm. It’s just that it’s an RL algorithm that involves a (possibly time- and situation-dependent) heuristic estimator of the expected value of a new random plan (a.k.a. the expected reward if you randomly tumble). If you’re way above that expected value, then keep doing whatever you’re doing; if you’re way below the threshold, re-rol... (read more)

[-]habryka9mo*1313

This reminds me of a conversation I had recently about whether the concept of "evil" is useful. I was arguing that I found "evil"/"corruption" helpful as a handle for a more model-free "move away from this kind of thing even if you can't predict how exactly it would be bad" relationship to a thing, which I found hard to express in a more consequentialist frames.

2tailcalled8mo

I feel like "evil" and "corruption" mean something different. Corruption is about selfish people exchanging their power within a system for favors (often outside the system) when they're not supposed to according to the rules of the system. For example a policeman taking bribes. It's something the creators/owners of the system should try to eliminate, but if the system itself is bad (e.g. Nazi Germany during the Holocaust), corruption might be something you sometimes ought to seek out instead of to avoid, like with Schindler saving his Jews. "Evil" I've in the past tended to take to take to refer to a sort of generic expression of badness (like you might call a sadistic sexual murderer evil, and you might call Hitler evil, and you might call plantation owners evil, but this has nothing to do with each other), but that was partly due to me naively believing that everyone is "trying to be good" in some sense. Like if I had to define evil, I would have defined it as "doing bad stuff for badness's sake, the inversion of good, though of course nobody actually is like that so it's only really used hyperbolically or for fictional characters as hyperstimuli". But after learning more about morality, there seem to be multiple things that can be called "evil": * Antinormativity (which admittedly is pretty adjacent to corruption, like if people are trying to stop corruption, then the corruption can use antinormativity to survive) * Coolness, i.e. countersignalling against goodness-hyperstimuli wielded by authorities, i.e. demonstrating an ability and desire to break the rules * People who hate great people cherry-picking unfortunate side-effects of great people's activities to make good people think that the great people are conspiring against good people and that they must fight the great people * Leaders who commit to stopping the above by selecting for people who do bad stuff to prove their loyalty to those leaders (think e.g. the Trump administration) I think "evil

[-]leogao8mo*120

i don't think this is unique to world models. you can also think of rewards as things you move towards or away from. this is compatible with translation/scaling-invariance because if you move towards everything but move towards X even more, then in the long run you will do more of X on net, because you only have so much probability mass to go around.

i have an alternative hypothesis for why positive and negative motivation feel distinct in humans.

although the expectation of the reward gradient doesn't change if you translate the reward, it hugely affects the variance of the gradient.^[1] in other words, if you always move towards everything, you will still eventually learn the right thing, but it will take a lot longer.

my hypothesis is that humans have some hard coded baseline for variance reduction. in the ancestral environment, the expectation of perceived reward was centered around where zero feels to be. our minds do try to adjust to changes in distribution (e.g hedonic adaptation), but it's not perfect, and so in the current world, our baseline may be suboptimal.

^{^}
Quick proof sketch (this is a very standard result in RL and is the motivation for advantage estimation, b

... (read more)

3Richard_Ngo8mo

I have a mental category of "results that are almost entirely irrelevant for realistically-computationally-bounded agents" (e.g. results related to AIXI), and my gut sense is that this seems like one such result.

2Garrett Baker8mo

I mean this situation is grounded & formal enough you can just go and implement the relevant RL algorithm and see if its relevant for that computationally bounded agent, right?

2leogao8mo

is this for a reason other than the variance thing I mention? I think the thing I mention is still important is because it means there is no fundamental difference between positive and negative motivation. I agree that if everything was different degrees of extreme bliss then the variance would be so high that you never learn anything in practice. but if you shift everything slightly such that some mildly unpleasant things are now mildly pleasant, I claim this will make learning a bit faster or slower but still converge to the same thing.

6Richard_Ngo8mo

Suppose you're in a setting where the world is so large that you will only ever experience a tiny fraction of it directly, and you have to figure out the rest via generalization. Then your argument doesn't hold up: shifting the mean might totally break your learning. But I claim that the real world is like this. So I am inherently skeptical of any result (like most convergence results) that rely on just trying approximately everything and gradually learning which to prefer and disprefer.

8leogao8mo

are you saying something like: you can't actually do more of everything except one thing, because you'll never do everything. so there's a lot of variance that comes from exploration that multiplies with your O(k2) variance from having a suboptimal zero point. so in practice your k needs to be very close to optimal. so my thing is true but not useful in practice. i feel people do empirically shift k quite a lot throughout life and it does seem to change how effectively they learn. if you're mildly depressed your k is slightly too low and you learn a little bit slower. if you're mildly manic your k is too high and you also learn a little bit slower. therapy, medications, and meditations shift k mildly.

[-]Vanessa Kosoy8moΩ5110

In (non-monotonic) infra-Bayesian physicalism, there is a vaguely similar asymmetry even though it's formalized via a loss function. Roughly speaking, the loss function expresses preferences over "which computations are running". This means that you can have a "positive" preference for a particular computation to run or a "negative" preference for a particular computation not to run^[1].

^{^}
There are also more complicated possibilities, such as "if P runs then I want Q to run but if P doesn't run then I rather that Q also doesn't run" or even preferences that are only expressible in terms of entanglement between computations.

8Kaj_Sotala9mo

Reminds me of @MalcolmOcean 's post on how awayness can't aim (except maybe in 1D worlds) since it can only move away from things, and aiming at a target requires going toward something.

5cubefox9mo

In Richard Jeffrey's utility theory there is actually a very natural distinction between positive and negative motivations/desires. A plausible axiom is U(⊤)=0 (the tautology has zero desirability: you already know it's true). Which implies with the main axiom[1] that the negation of any proposition with positive utility has negative utility, and vice versa. Which is intuitive: If something is good, its negation is bad, and the other way round. In particular, if U(X)=U(¬X) (indifference between X and ¬X), then U(X)=U(¬X)=0. More generally, U(¬X)=−(P(X)/P(¬X))U(X). Which means that positive and negative utility of a proposition and it's negation are scaled according to their relative odds. For example, while your lottery ticket winning the jackpot is obviously very good (large positive utility), having a losing ticket is clearly not very bad (small negative utility). Why? Because losing the lottery is very likely, far more likely than winning. Which means losing was already "priced in" to a large degree. If you learned that you indeed lost, that wouldn't be a big update, so the "news value" is negative but not large in magnitude. Which means this utility theory has a zero point. Utility functions are therefore not invariant under adding an arbitrary constant. So the theory actually allows you to say X is "twice as good" as Y, "three times as bad", "much better" etc. It's a ratio scale. ---------------------------------------- 1. If P(X∧Y)=0 and P(X∨Y)≠0 then U(X∨Y)=P(X)U(X)+P(Y)U(Y)P(X)+P(Y). ↩︎

[-]Richard_Ngo2yΩ19433

Five clusters of alignment researchers

Very broadly speaking, alignment researchers seem to fall into five different clusters when it comes to thinking about AI risk:

MIRI cluster. Think that P(doom) is very high, based on intuitions about instrumental convergence, deceptive alignment, etc. Does work that's very different from mainstream ML. Central members: Eliezer Yudkowsky, Nate Soares.
Structural risk cluster. Think that doom is more likely than not, but not for the same reasons as the MIRI cluster. Instead, this cluster focuses on systemic risks, multi-agent alignment, selective forces outside gradient descent, etc. Often work that's fairly continuous with mainstream ML, but willing to be unusually speculative by the standards of the field. Central members: Dan Hendrycks, David Krueger, Andrew Critch.
Constellation cluster. More optimistic than either of the previous two clusters. Focuses more on risk from power-seeking AI than the structural risk cluster, but does work that is more speculative or conceptually-oriented than mainstream ML. Central members: Paul Christiano, Buck Shlegeris, Holden Karnofsky. (Named after Constellation coworking space.)
Prosaic cluster. Focuses on empi

... (read more)

[-]Richard_Ngo4mo*38-44

Crossposted from Twitter:

This year I’ve been thinking a lot about how the western world got so dysfunctional. Here’s my rough, best-guess story:

1. WW2 gave rise to a strong taboo against ethnonationalism. While perhaps at first this taboo was valuable, over time it also contaminated discussions of race differences, nationalism, and even IQ itself, to the point where even truths that seemed totally obvious to WW2-era people also became taboo. There’s no mechanism for subsequent generations to create common knowledge that certain facts are true but usefully taboo—they simply act as if these facts are false, which leads to arbitrarily bad policies (e.g. killing meritocratic hiring processes like IQ tests).

2. However, these taboos would gradually have lost power if the west (and the US in particular) had maintained impartial rule of law and constitutional freedoms. Instead, politicization of the bureaucracy and judiciary allowed them to spread. This was enabled by the “managerial revolution” under which govt bureaucracy massively expanded in scope and powers. Partly this was a justifiable response to the increasing complexity of the world (and various kinds of incompetence and nepotism... (read more)

[-]lc4mo*6763

This class gains status by signaling commitment to luxury beliefs. Since more absurd beliefs are more costly-to-fake signals, the resulting ideology is actively perverse (i.e. supports whatever is least aligned with their stated core values, like Hamas).

Without commenting on your broader point, I think you believe elites support Hamas because the conservative Twitter feed is presenting you/your circle tailored ragebait, not because elites invert their own utility functions.

1Richard_Ngo4mo

That's a mechanism by which I might overestimate the support for Hamas. But the thing I'm trying to explain is the overall alignment between leftists and Hamas, which is not just a twitter bubble thing (e.g. see university encampments). More generally, leftists profess many values which are upheld the most by western civilization (e.g. support for sexual freedom, women's rights, anti-racism, etc). But then in conflicts they often side specifically against western civilization. This seems like a straightforward example of pessimization.

[-]lc4mo*3725

More generally, leftists profess many values which are upheld the most by western civilization (e.g. support for sexual freedom, women's rights, anti-racism, etc). But then in conflicts they often side specifically against western civilization. This seems like a straightforward example of pessimization.

Not at all. The trend is that in any given context, American leftists tend to support the 'weaker' group against the stronger group, regardless of the merits of the individual cases. They have a world model that says that that most social problems come from "big" people hurting "little" people, and believe the focus of their politics should be remedying this. In the case of Israel-Palestine, Israel and the United States are much more militarily and economically powerful than Gaza, so ceteris paribus^[1] they side with Gaza, just as they side with women, ethnic minorities, the poor, etc. You may disagree with this behavior but it's fairly consistent.

By contrast, the argument that Israel is a bastion of western values and therefore leftists should support its war against a smaller neighbor is kind of abstract. The immediate outcome of Israel winning the war is just that Israe... (read more)

2Richard_Ngo1mo

A few responses: * As per my post on underdog bias, the question of which group is actually weaker and which group is stronger is often a pretty subjective call. I even discuss in the post the example of Israel, where you could see it as the "stronger" group (vs Palestine in particular) or the "weaker" group (vs all the Muslim countries surrounding it). * There are plenty of cases where leftists support the stronger group against the weaker group—most notably Soviet and Chinese repression of dissidents and minorities. E.g. it took Solzhenitsyn publishing Gulag Archipelago to finally get leftists (even fairly "mainstream" leftists) to stop lionizing the USSR. * Even insofar as leftists tend to support the weaker group, there are almost no cases where they do so as strongly as in Israel vs Palestine. So there's still something important to be explained here even accepting your claims.

[-][anonymous]4mo158

This characterizes leftists sufficiently dishonestly that I think you've gotten mindkilled by politics. As people keep removing my (entirely accurate and if anything understated) soldier-mindset reacts, I will strong-downvote instead :)

2Richard_Ngo1mo

Wanted to revisit this because it seemed like one of the points where people most strongly disagreed with me. I'm trying to figure out a crux here. One might be something like: how widespread were celebrations of the 10/7 attacks amongst prominent leftists (and especially the student groups that later organized encampments)? I could imagine updating that there were only a handful of cases that disproportionately blew up, which would make me take back or at least caveat the "supports Hamas" thing. If you on the other hand found that there were many cases where prominent leftists and encampment organizers actively celebrated the 10/7 attacks, would you (@dirk or @lc) then agree that "supports Hamas" is a reasonable summary?

4speck14471mo

I am not one of the tagged people but I certainly would not so agree. One reason I would not so agree is because I have talked to leftist people (prominence debatable) who celebrated the 10/7 attacks, and when I asked them whether they support Hamas, they were coherently able to answer "no, but I support armed resistance against Israel and don't generally condemn actions that fall in that category, even when I don't approve of or condone the group organizing those actions generally." One way to know what people believe and support is to ask them. (Of course, I don't think this is a morally acceptable position either, and conversation ensued! But it's clearly not "supporting Hamas" in any sense that can support your original claims.) My social circles also include many leftists, including student organizers and somewhat well-known online figures, so I separately suspect that you're vastly overestimating the proportion of self-identified leftists who celebrated the attacks in any meaningful sense, but that's probably not the crux here.

5the gears to ascension4mo

I think you're confusing levels of behavior, here. It may be that their net effect is to support Hamas, but that they don't intend this, and that their actual intent does not make the mistake you're (at least if I read correctly) attributing to intent here. I do think they end up on net having intentions that make them vulnerable to being manipulated by those in power in adversary groups, since their intent is to support the weak among any group. In particular, they typically are thinking in different group selectors than you seem to assign here; of the people I'd guess you mean, in my interactions with them, they don't seem to profess support for Hamas, and in fact explicitly say they'd like to support palestinians without supporting hamas. But, I strongly agree that leftists tend to pessimize their values on net, and that that is one of the biggest issues with leftist approaches to the world. So whether we disagree depends on what level we're looking at. Your reasoning seems like it would be improved by finding people who seem to exhibit the belief you suspect exists and want to model, and interviewing them without letting your opinions leak, to try to get a map of their actual opinions; before making this comment I spent some time brainstorming to find a way to do that fast, and didn't come up with one, so perhaps it makes sense to not take this suggestion, but I maintain that it would be useful if practical. As a stopgap, going to the main locations that a group discuss things online can be useful, keeping in mind that then one will have uncertainty like the kind I have about you: you seem overall to be a good person as far as I've been able to tell so far, but you interact enough with people whose personal policies in practice support things that seem to me like they're at imminent risk of causing human-induced catastrophic outcomes for many of the humans in the world that it's unclear to me what your intentions are.

[-]johnswentworth4mo2918

Without commenting on the specifics, I agree with a lot of the gestalt here as a description of how things evolved historically, but I think that's not really the right lens to understand the problem.

My current best one-sentence understanding: the richer humans get, the more social reality can diverge from physical reality, and therefore the more resources can be captured by parasitic egregores/memes/institutions/ideologies/interest-groups/etc. Physical reality provides constraints and feedback which limit the propagation of such parasites, but wealth makes the constraints less binding and therefore makes the feedback weaker.

[-]Linch4mo127

The main reason I disagree with both this comment and the OP is that you both have the underlying assumption that we are in a nadir (local nadir?) of connectedness-with-reality, whereas from my read of history I see no evidence of this, and indeed plenty of evidence against.

People used to be confused about all sorts of things, including, but not limited to, the supernatural, the causes of disease, causality itself, the capabilities of women, whether children can have conscious experiences, and so forth.

I think we've gotten more reasonable about almost everything, with a few minor exceptions that people seem to like highlighting (I assume in part because they're so rare).

The past is a foreign place, and mostly not a pleasant one.

7johnswentworth4mo

I totally buy that peoples' verbal models aren't at a local nadir of connectedness-to-reality. The thing which seems increasingly disconnected from reality is more like metis, peoples' day-to-day behavior and intuitive knowledge, institutional knowledge and skills, personal identity and goals, that sort of thing. I'm notably not thinking here primarily about examples like e.g. heritability of IQ becoming politicized; that's a verbal model, and I do think that verbal models have mostly become more reasonable modulo a few exceptions which people highlight.

2Richard_Ngo4mo

I used to agree with your understanding but I am now more skeptical. For example, here's a story that says the opposite: I'm not saying my story is true, but it does highlight that the load-bearing question is actually something like "how does the offense-defense balance against parasitic egregores scale with wealth?" Why don't we live in a world where wealth can buy a society defenses against such egregores? Or maybe we do live in such a world, and we are just failing to buy those defenses. That seems like a really dumb situation to be in, but I think my post is broadly describing how it might arise.

9johnswentworth4mo

I would point to the non-experts can't distinguish true from fake experts problem. That does seem to be a central phenomenon which most parasitic egregores exploit. More generally, as wealth becomes more abundant (and therefore lots of constraints become more slack), inability to get grounded feedback becomes a more taut constraint. That said... do you remember any particular evidence or argument which led you toward the story at top of thread (as opposed to away from your previous understanding)?

[-]Wei Dai4mo*280

Can you explain more your affinity for virtue ethics, e.g., was there a golden age in history, that you read about, where a group of people ran on virtue ethics and it worked out really well? I'm trying to understand why you seem to like it a lot more than I do.

Re government debt, I think that is actually driven more by increasing demand for a "risk-free" asset, with the supply going up more or less passively (what politician is going to refuse to increase debt and spending, as long as people are willing to keep buying it at a low interest rate). And from this perspective it's not really a problem except for everyone getting used to the higher spending when some of the processes increasing the demand for government debt might only be temporary.

AI written explanation of how financialization causes increased demand for government debt

- Financialization isn't a vague blob; it's a set of specific, concrete processes, each of which acts like a powerful vacuum cleaner sucking up government debt.
  Let's trace four of the most important mechanisms in detail.
  1. The Derivatives Market: The Collateral Multiplier
  Derivatives (options, futures, swaps) are essentially financial side-bets on the movem

... (read more)

[-]Eli Tyre4mo2714

...if the west (and the US in particular) had maintained impartial rule of law and constitutional freedoms.

The US did not have impartial rule of law in the late 1800s and early 1900s. Notably, black Americans in the south were regularly impressed into forced labor, often for the rest of their lives, on the basis of flimsy or even non-existent legal pretext.

(A representative but concocted example: the local sheirf arrests a black man who's walking through town on charges of "vagrancy". The man is found guilty and sentenced to hard labor. The sheriff sells the "contact" for hard labor to a local industrialist who owns a mine. (The sherif and the industrialist are buddies, and have done versions of this deal many times before). The man is set to work the mine for the period of his sentence. When he's near the end of his sentence, he's accused of some minor infraction as a pretext to add more years to his sentence (if anyone bothers to keep track of when his sentence is served at all.)

What makes you think that impartial rule of law decayed since WWII instead of generally (though not evenly) improving?

4Richard_Ngo1mo

That's a fair point; however, I don't think it undermines my overall claims very much. I think the lack of rule of law for black Americans was bad in a comparable way to how the lack of rule of law for various European colonies was bad. That is, while it was bad for the people who didn't get rule of law, they were a separate enough category that this mostly didn't "leak into" undermining the legal mechanisms that helped their societies become productive and functional in the first place.

[-]Eli Tyre1mo1415

That is, while it was bad for the people who didn't get rule of law, they were a separate enough category that this mostly didn't "leak into" undermining the legal mechanisms that helped their societies become productive and functional in the first place.

I'm speaking speculatively here, but I don't know that it didn't leak out and undermine the mechanism that supported productive and functional societies. The sophisticated SJW in me suggests that this is part of what caused the eventual (though not yet complete) erosion of those mechanisms.

It seems like if you have "rule of law" that isn't evenly distributed, actually what you have is collusion by one class of people to maintain a set of privileges at the expense of another class of people, where one of the privileges is a sand-boxed set of norms that govern dealings within the privileged class, but with the pretense that the norms are universal.

This kind of pretense seems like it could be corrosive: people can see that the norms that society proclaims as universal actually aren't. This reinforces a a sense that the norms aren't real at all (or at least) a justified sense that the ideals that underly those norms are mostly rational... (read more)

[-]Shiva's Right Foot4mo134

This makes no mention of the repeal of the fairness doctrine nor the shift in financial model for major newspapers. The 1987 abolishment of the Fairness Doctrine led very directly to Rush Limbaugh gaining a national political audience.

The fairness doctrine of the United States Federal Communications Commission (FCC), introduced in 1949, was a policy that required the holders of broadcast licenses both to present controversial issues of public importance and to do so in a manner that fairly reflected differing viewpoints.[1] In 1987, the FCC abolished the fairness doctrine,[2] prompting some to urge its reintroduction through either Commission policy or congressional legislation.[3] The FCC removed the rule that implemented the policy from the Federal Register in August 2011.[4]

Rush Limbaugh used to be a regular music DJ in the 1970s. His political talk show was distributed nationally in 1988, soon after the cessation of Fairness Doctrine support by the Reagan administration. Carl McIntire was the Rush Limbaugh (maybe a little more Glenn Beck) of the 1960s and had his radio talk show shut down by the Fairness Doctrine after years of litigation. The legal challenges the Fairness ... (read more)

[-]testingthewaters4mo135

To be honest, this post approaches a level of disjointedness from reality-as-I-understand-it that I fear I cannot accurately respond to it in a way that would satisfy the author.^[1] However, if I don't give some response, I suspect that it will join a growing cultural current within LessWrong which involves an intellectualised rationalisation of ethnonationalism, hygiene-oriented eugenics, oligarchy^[2], and implicit or explicit support of political action and political violence for these ends. If this current becomes normalised (and it has, frankly, always been present in the rationalist sphere), it makes TESCREAL-style categorisations of the rationalist/EA/AI safety intellectual sphere bascially correct. Therefore, I want to commit to voicing my objection to this line of argument where I see it. There are parts of me that feel strongly against doing this. There is no glory to be gained in political arguments online. However, there is the possibility of avoiding shame, which as a virtue ethicist I'm sure the author will understand.

On the question of inherent racial or ethnic attributes, I suspect that I will not be able to settle the argument "on the facts", given how long the... (read more)

[-]Eli Tyre4mo*110

If this current becomes normalised (and it has, frankly, always been present in the rationalist sphere), it makes TESCREAL-style categorisations of the rationalist/EA/AI safety intellectual sphere bascially correct.

What? It seems like TESCREAL is clearly a natural cluster (modulo I don't know any examples of the "C"), and whether takes like this are pervasive amongst rationalists doesn't bare on whether it's a good categorization?

8Eli Tyre4mo

I'm not sure how this is a response to the OP. It sounds basically right to me (and I imagine that Richard would agree with it as well, though he can speak for himself), but it seems like almost entirely a non sequitur to the claims made? This text doesn't even mention ethnic segregation as a solution? It does promote virtue ethics as an alternative moral frame to utilitarianism, but it says nothing about "enforced[ing] by inter-group mutual distrust". I don't doubt that you're pushing back against a real cultural force, but it doesn't look to me to actually be represented in this short-form, except (at best) as an implication.

8Richard_Ngo4mo

Might reply to the rest later but just to respond to what you call "the most obvious example": consider a company which has a difficult time evaluating how well its employees are performing (i.e. most of them). Some employees will work hard even when they won't directly be rewarded for that, because they consider it virtuous to do so. However, if you then add to their team a bunch of other people who are rewarded for slacking off, the hard-working employees may become demotivated and feel like they're chumps for even trying to be virtuous. The extent to which modern governments hand out money causes a similar effect across western societies (edited: for example, if many people around you are receiving welfare, then working hard yourself is less motivating). They would not be as able to do this as much if their currencies were still on the gold standard, because it would be more obvious that they are insolvent.

8testingthewaters4mo

I'm afraid, if you're actually trying to advance an argument, you're going to need to be slightly more specific. Which government, in which society? Perhaps you will say, "all of the western governments". Very well. Then which gold standard? I will remind you that "western societies" including Germany, the UK, and the US exited and entered the "gold standard", or swapped out various forms of "gold standard", all throughout the 20th century. Was Britain in 1931 or Germany in 1914 lacking in valor? I'm by no means an economic historian, and this is just wikipedia talking. And even when countries obeyed the "gold standard", policies varied strongly between regions, including how much currency could be redeemed for how much gold, whether citizens could redeem gold directly, the status of gold as a commodity, the trade of gold between countries etc. In what combination of policies can we find your notion of "honest work for honest pay"? Does virtue's just reward include the ability to hold gold sovereigns as private property to be redeemed at any time at a national bank, or to exchange gold bullion across national borders? If so, at what exchange rates? And what does "insolvent" mean? After world war II the world's currencies were arranged such that the US dollar was the world's reserve currency as part of the Bretton Woods agreement. The US obtained an exorbitant privilege, in that it could print paper and get money while other countries needed to produce goods or services. Was this the point at which the US became, as you say, "insolvent"? Yet it was the paramount superpower at the time, and the value of its currency was (supposedly) backed by gold - until it decided to break from the gold standard under Nixon and still more or less retained that privilege. But was America's massive spending during the post world war II period and the Cold War (all powered directly or indirectly by this exorbitant privilege) a sign of degeneracy and decay? Surely not. Okay, let

[-]Richard_Ngo4mo10-1

Thanks for the extensive comment. I'm not sure it's productive to debate this much on the object level. The main thing I want to highlight is that this is a very good example of how the taboo that I discussed above operates.

On most issues, people (and especially LWers) are generally open to thinking about the benefits and costs of each stance, since tradeoffs are real.

However, in the case of ethnonationalism, even discussing the taboo on it (without explicitly advocating for it) was enough to trigger a kind of zero-tolerance attitude in your comment.

This is all the more striking because the main historical opponent of ethnonationalist regimes was globalist communism, which also led to large-scale atrocities. Yet when people defend a "socialist" or "egalitarian" cluster of ideas, that doesn't lead to anywhere near this level of visceral response.

My main bid here is for readers to notice that there is a striking asymmetry in how we think about and discuss 20th century history, which is best explained via the thing I hypothesized above: a strong taboo on ethnonationalism in the wake of WW2, which has then distorted our ability to think about many other issues.

[-]testingthewaters4mo1318

If you press them too closely, they will abruptly fall silent, loftily indicating by some phrase that the time for argument is past.

Jean-Paul Sartre

I sincerely believe that people will get hurt if these ideas return to society at large, Richard. Please don't do this.

[-]Wei Dai4mo226

It seems to me that Richard isn't trying to bring back ethnonationalism, or even trying to "add just that touch of ethnic pride back into the meme pool", but just trying to diagnose "how the western world got so dysfunctional". If ethnonationalism and the taboo against ethnonationalism are both bad (as an ethnic minority I'm personally pretty scared of the former), then maybe we should get rid of the taboo and defend against ethnonationalism by other means, similar how there is little to no taboo against communism^[1] but it hasn't come close to taking power or reapproaching its historical high water mark in the west.

^{^}
If you doubt this, there's an advisor to my local school district who is a self-avowed Marxist and professor of education at the state university, and writes book reviews like this one:
«For decades the educational Left and critical pedagogues have run away from Marxism, socialism, and communism, all too often based on faulty understandings and falling prey to the deep-seated anti-communism in the academy. In History and Education Curry Stephenson Malott pushes back against this trend by offering us deeply Marxist thinking about the circulation of capital, socialis

... (read more)

7Richard_Ngo4mo

I like this comment. For the sake of transparency, while in this post I'm mostly trying to identify a diagnosis, in the longer term I expect to try to do political advocacy as well. And it's reasonable to expect that people like me who are willing to break the taboo for the purposes of diagnosis will be more sympathetic to ethnonationalism in their advocacy than people who aren't. For example, I've previously argued on twitter that South Africa should have split into two roughly-ethnonationalist states in the 90s, instead of doing what they actually did. However, I expect that the best ways of fixing western countries won't involve very much ethnonationalism by historical standards, because it's a very blunt tool. Also, I suspect that breaking the taboo now will actually lead to less ethnonationalism in the long term. For example, even a little bit more ethnonationalism would plausibly have made European immigration policies much less insane over the last few decades, which would then have prevented a lot of the political polarization we're seeing today.

3cdt4mo

I find this idea particularly fraught. I already find it somewhat difficult to engage on this site due to the contentious theories some members hold, and I echo testingthewater's warning against the trap of reopening these old controversies. You're trying to thread a really fine needle between "meaningfully advocate change" and "open all possible debates" that I don't think is feasible. The site is currently watching a major push from Yudkowsky and Soares' book launch towards a broad coalition for an AI pause. It really only takes a couple major incidents of connecting the idea to ethnonationalism, scientific racism and/or dictatorship for the targets of your advocacy to turn away. I'm not going to suggest you stay on-message (lw is way too "truth-seeking" for that to reach anyone), but you should carefully consider the ways in which your future goals conflict.

[-]the gears to ascension4mo*112

I could imagine a world where discussing it is not something I would see as critically unwelcome, but I would still want to maintain part of the mechanism that implements the taboo, which is why I made the comments I did: being able to come to agreement that there are outcomes which are good and bad, such as the previous famous outcomes of ethnonationalism. I suspect that I've interpreted you to be saying that those previous outcomes were bad ("While perhaps at first this taboo was valuable") but it's not as reliable as the taboo implementation I carry would prefer.

If your view is that virtues are required to be independent of outcomes, in that eg a virtue whose adoption would predictably lead to mass death [edit to finish phrase: would still be a virtue if virtuous by its own merits], then I don't think they can be truly called virtuous; but if your view is that we can discuss some features of outcomes that make an outcome good or bad, then I would want to discuss what virtues can lead us towards better outcomes by that standard. virtue utilitarianism, if you will, rather than rule utilitarianism^[1]; the way I heard virtues described was that a virtue is a local aspiration, a non-... (read more)

8Richard_Ngo4mo

This is a thoughtful comment, I appreciate it, and I'll reply when I have more time (hopefully in a few days).

6lc4mo

I seriously doubt Richard recommends any of the regimes/interventions you actually argue against here.

0Viliam4mo

I think the problem is that a part of "respecting" people is letting them choose things for themselves, and in a democratic society also letting them choose things for others. I admit I do have a problem respecting many people in this specific way. Not sure what to do about it though.

-5Big Tony4mo

[-]Viliam4mo122

Human perception of society has some paradoxes. Consider freedom of speech: In countries that generally have freedom of speech, many people complain about all kinds of injustice, censorship, etc. In countries that have no freedom of speech, everyone is quiet, and when asked explicitly, says: everything is great. Therefore, naive observers often conclude that the former countries have less freedom of speech than the latter, judging by the number of complaints about censorship.

I believe there is a similar effect with meritocracy/equality/etc. Imagine a perfectly unfair feudal society where unless you are born as a member of aristocracy, you are screwed; your talents and hard work will make absolutely no difference. Ironically, many people will believe that this society is fair, that the aristocrats are chosen by God for being better. If the poor kids are never given an opportunity to learn, everyone may believe, based on what they observe, that the poor kids are completely unable to learn. This is what all their priests would teach.

Then comes a revolution, and people find out that the aristocrats are often stupid, and that if you give free education to the poor kids, many of them tur... (read more)

9Jonas Hallgren4mo

I was reflecting on some of the takes here for a bit and if I imagine a blind gradient descent in this direction, I imagine quite a lot of potential reality distortion fields due to various of the underlying dynamics involved with holding this position. So the one thing I wanted to ask was that if you have any sort of reset mechanism here? Like what is the schelling point before the slippery slope? What is the specific action pattern you would take if you got too far? Or do you trust future you enough in order to ensure that it won't happen?

3Richard_Ngo4mo

Good question. One answer is that my reset mechanisms involve cultivating empathy, and replacing fear with positive motivation. If I notice myself being too unempathetic or too fear-driven, that's worrying. But another answer is just that, unfortunately, the reality distortion fields are everywhere—and in many ways more prevalent in "mainstream" positions (as discussed in my post). Being more mainstream does get you "safety in numbers"—i.e. it's harder for you to catalyze big things, for better or worse. But the cost is that you end up in groupthink.

5Eli Tyre4mo

Can you give some examples? Big government debts seem like the kind of thing that can fail catastrophically at some point, so they represent an important civilizational risk. But how are they distorting the economy? Do you just mean that governments spend much more than they raise in taxes?

2wassname3mo

This is a theory often referred to as the Cantillian effect Richard Cantillon observed that the original recipients of new money enjoy higher standards of living at the expense of later recipients. In colloquial terms, the closer you stand to the source of money creation, the wealthier you become. When governments run large deficits that get monetized by central banks, this creates new money that flows first to government and financial sectors before reaching the broader economy. This distorts resource allocation because entities closer to the money source can bid up assets and resources before prices adjust throughout the system.

5Garrett Baker4mo

From a rhetorical perspective, I think it was wrong to lead with the ethnonationalism stuff, since it seems to have confused testingthewaters, and likely many more.

4Douglas_Knight4mo

The points about IQ seem parochially American, not applicable to the rest of the West. But aren't you British, not American? Does this really seem so central in Britain?

2the gears to ascension4mo

I find it hard to tell what you see as good or bad, among these things, especially around what currently is, vs what you would see in the world. I of course can read what you've explicitly stated as good, but the way you've written this puts it nearby in linguistic vector space to things which are controversial at best, but are intertwined with processes that I would expect to cause seriously bad things. I would appreciate if you could find a way to more clearly signal what outcomes you are inclined towards, to the degree that avoiding accidental pessimization permits you to do so, separate from how to get there. Because as is, it looks to me like you're seeing the world through a ... very strange lens, one that sees real problems but magnifies and distorts them in ways I find unfamiliar and confusing, and which from my understanding seems to blur things together; and I am unclear whether we have common ground, whether the things I see as good in the world are things you see as good, whether things I see as major evils of society are things you do. You've explicitly said that some things which attempt to do things in the world under banners of things one might think are good have instead done ill, and things that one might think do ill can be good. In your talk I didn't have an opportunity to properly probe your models, but it seemed to me that your label of evil was, at a minimum, not the core etiology of evil, and I worry that if what you seek as good really is simply the other quadrant, that you're making a subtle mistake. If you can look at what you wrote, and why I would write this, I think you will see that you have already described part of my reaction partially; but I think you have misunderstood my reaction, and I don't feel that it is correct for me to fully clarify unless you can take a step of clarifying what it is you seek... virtues, I suppose, but I would want to know what those are. I hope you can forgive my intentional vagueness. edit: hash of t

2Vladimir_Nesov4mo

As git hashsums are short and tangible True Names of abstract git objects, there are abstract properties of behaviors of things in the world that are in principle as concrete and tangible as coal or silver. The economy uses abstract goods to produce new abstract goods. Consequentialism and utility functions or policies could in principle be about virtues and integrity as about hamburgers, but hamburgers are more legible and easier to administer. So I think the crux is relative legibility rather than methods, the same methods that should in principle work break down for practical reasons that have nothing to do with applicability of the methods in principle.

2Richard_Ngo4mo

Here's one concrete way in which this isn't true: one common simplifying assumption in economics is that goods are homogeneous, and therefore that you're indifferent about who to buy from. However, virtuous behavior involves rewarding people you think are more virtuous (e.g. by preferentially buying things from them). In other words, economics is about how agents interact with each other via exchanging goods and services, while virtues are about how agents interact with each other more generally.

4Garrett Baker4mo

I don't see why using the word "virtue" magically solves the hardness of that math (the reasons why such assumptions are simplifying). Maybe that is what your promised formalizations are about, but I'm also kinda skeptical that reputation and consumers' preference to trade with some groups over others is a subject that no economist could give reasonable models about.

4Vladimir_Nesov4mo

It seems very reasonable to be indifferent about who ends up incentivising or exhibiting virtuous behaviors you care about (to be more prevalent in the world or your community), and the incentives don't need to come in the form of personal action about who to buy other things from. If virtues and other abstract properties of behaviors are treated as particular examples of goods and services, then economics could discuss how agents obtain the presence of virtues or patterns of interaction between people, by paying businesses that specialize in manufacturing their presence in the world. This does need a scalable enough business to merit the term that can produce marginal virtue and patterns of interactions by employing the existing economy, rearranging the physical world in a way that results in greater presence of these abstract goods in it, wielding fiat currency to move other goods and labor in the world to make this happen by paying other businesses and people who specialize in those goods and labor. Some of these other goods instrumentally effected through other businesses could themselves be virtues or patterns of interaction. Building economic engines that scale is very hard (successful startups reward founders and investors), doing this with illegible abstract goods is borderline impossible, but this is not a fundamentally different kind of activity.

-8Cole Wyeth4mo

[-]Richard_Ngo2y*341

(COI note: I work at OpenAI. These are my personal views, though.)

My quick take on the "AI pause debate", framed in terms of two scenarios for how the AI safety community might evolve over the coming years:

AI safety becomes the single community that's the most knowledgeable about cutting-edge ML systems. The smartest up-and-coming ML researchers find themselves constantly coming to AI safety spaces, because that's the place to go if you want to nerd out about the models. It feels like the early days of hacker culture. There's a constant flow of ideas and brainstorming in those spaces; the core alignment ideas are standard background knowledge for everyone there. There are hackathons where people build fun demos, and people figuring out ways of using AI to augment their research. Constant interactions with the models allows people to gain really good hands-on intuitions about how they work, which they leverage into doing great research that helps us actually understand them better. When the public ends up demanding regulation, there's a large pool of competent people who are broadly reasonable about the risks, and can slot into the relevant institutions and make them work well.
AI sa

... (read more)

9Raemon2y

FYI I think this is worth fleshing out into a top level post (esp. given that it's 'Pause Debate' week). I'm not actually sure it needs much fleshing out. I think the main bit here that feels unjustified, or insufficiently-justified for the strength of the claim, is:

2Richard_Ngo2y

Have edited slightly to clarify that it was "leading experts" who dramatically underestimated it. I'm not really sure what else to say, though...

4Daniel Kokotajlo2y

I think I basically agree with you Richard about the risks of falling into scenario 2, and think this is a wise comment, but I also think you are strawmanning the reason for the change -- it's not that people have come to think that making technical progress is hopeless (or even harder than it used to be!) it's rather that people have come to have shorter timelines, and so the probability that sufficient technical progress will be made in time has gone down, and the usefulness of calling for a pause has gone up. (e.g. if you think AGI is 15 years away, then pausing now is plausibly useless or even harmful.) That's my theory at any rate. And it's sorta what I think, I think. Oh, also, I think that the counterproductive rot in environmentalism took at least 5 years to build up, probably, and I'm hopeful that therefore even if we are on a path to rot, it'll take too long for the rot to build up to matter. But this is just a guess about the growth rate of rot which is informed by anecdotes like the stories about what happened to your friends, and over the coming months and years more data will be collected to better calibrate my guess about the rate of rot.

2mako yass2y

I appreciated seeing the caricature of 1 presented at least, as a dream, it feels attainable, or like it might have been in some other timeline, but perhaps in ours too, for all I know.

[-]Richard_Ngo1moΩ17330

Error-correcting codes work by running some algorithm to decode potentially-corrupted data. But what if the algorithm might also have been corrupted? One approach to dealing with this is triple modular redundancy, in which three copies of the algorithm each do the computation and take the majority vote on what the output should be. But this still creates a single point of failure—the part where the majority voting is implemented. Maybe this is fine if the corruption is random, because the voting algorithm can constitute a very small proportion of the total code. But I'm most interested in the case where the corruption happens adversarially—where the adversary would home in on the voting algorithm as the key thing to corrupt.

After a quick search, I can't find much work on this specific question. But I want to speculate on what such an "error-correcting algorithm" might look like. The idea of running many copies of it in parallel seems solid, so that it's hard to corrupt a majority at once. But there can't be a single voting algorithm (or any other kind of "overseer") between those copies and the output channel, because that overseer might itself be corrupted. Instead, you need the m... (read more)

[-]Carl Feynman1moΩ123813

I have some experience in the design of systems designed for high reliability and resistance to adversaries. I feel like I’ve seen this kind of thinking before.

Your current line of thinking is at a stage I would call “pretheoretical noodling around.” I don’t mean any disrespect; all design has to go through this stage. But you’re not going to find any good references, or come to any conclusions, if you stay at this stage. A next step is to settle on a model of what you want to get done, and what capabilities the adversaries have. You need some bounds on the adversaries; otherwise nothing can work. And of course you need some bounds on what the system does, and how reliably. Once you’ve got this, you can either figure out how to do it, or prove that it can’t be done.

For example there are ways of designing hardware which is reliable on the assumption that at most N transistors are corrupt.

The problem of coming to agreement between a number of actors, some of whom are corrupt, is known as the Byzantine generals problem. It is well studied, and you may find it interesting.

I’m also interested in this topic, and I look forward to seeing where this line of thinking takes you.

2Richard_Ngo1mo

Perhaps. The issue here is that I'm not so interested in any specific goal, but rather in facilitating emergent complexity. One analogy here is designing Conway's game of life: I expect that it wasn't a process of "pick the rules you want, then see what results from those" but also in part "pick what results you want, and then see what rules lead to that". Re the Byzantine generals problem, see my reply to niplav below:

2Q Home1mo

I think he already came to some conclusions and you already gave some good references (which support some of the conclusions). Do those methods have names or address problems which have names (like the Byzantine generals problem)?

[-]Thoroughly Typed1mo184

This makes me think of "radiation-hardened" quines, quines that output themselves even after deleting any one character (or three):
https://github.com/mame/radiation-hardened-quine
https://codegolf.stackexchange.com/a/100785

I'm not familiar with the method used, but maybe it can be adapted to the case you are imagining? Maybe for simplicity only for the voting part of a redundant system.

5Richard_Ngo1mo

This seems very relevant, thank you. Will check it out.

[-]niplav1mo160

IIRC in Byzantine fault tolerance the assumption usually is that cryptographic primitives can be trusted and for $3 f + 1$ at most $f$ nodes are corrupted. The corrupted nodes can output anything. E.g. the Dolev-Strong protocol only assumes (1) a known number of nodes, (2) a public key infrastructure and (3) synchronized communication. But I guess your thought is around someone corrupting a specific part of all nodes? Dolev-Strong assumes you run more rounds than there are corrupted nodes.

1Richard_Ngo1mo

No, I'm happy to stick with the standard assumption of limited amounts of corruption. However, I believe (please correct me if I'm wrong) that Byzantine fault tolerance mostly thinks about cases where the nodes give separate outputs—e.g. in the Byzantine generals problem, the "output" of each node is whether it attacks or retreats. But I'm interested in cases where the nodes need to end up producing a "synthesis" output—i.e. there's a single output channel under joint control.

[-]Jesper L.1mo122

After a quick search, I can't find much work on this specific question. But I want to speculate on what such an "error-correcting algorithm" might look like.

Did you look in places discussing ledger tech and proof of stake? If you are curious to see what's been discussed, looking into that might be worth it. Information theory will mostly just focus on efficiency and practical trade-offs.

6mako yass1mo

I'd expect the answer to not be apparent to an outsider, reading the literature, but I'd expect people who are good at designing those sorts of systems to be able to give you the answer quite easily if you ask.

8Dalcy1mo

Not exactly about adversarial error correction, but: there is a construction (Çapuni & Gács 2021) of a (class of) universal 1-tape (!!) Turing machine that can perform arbitrarily long computation subject to random noise in the per-step action. Despite the non-adversarial noise model, naive majority error correction (or at least their construction of it) only fixes bounded & local error bursts - meaning it doesn't work in the general case, because even though majority vote reduces error probability, the effective error rate is still positive, so something almost surely goes wrong (eg error burst of size greater than what majority vote can handle) as T→∞. Their construction, in fact, looks like a hierarchy of simulated turing machines where the higher-level TM is simulated by a level below it but at a bigger tape scale, such that it can resist larger error bursts - and the overall construction looks like "moving" the "live simulation" of the actual program that we want to execute up the hierarchy over time to coarser and more reliable levels.

2Richard_Ngo1mo

That is very cool, thank you! Will check it out.

4Daniel C1mo

For this we need a mechanism such that the maintenance of the mechanism is a schelling point. Specifically, the mechanism at T+1 should reward agents for actions at time T that reinforce the mechanism itself (in particular the actions are distributed). The incentive raises the probability of the mechanism being actualized at T+1, which in turn raises the "weight" of the reward offered by the mechanism at T+1, creating a self-fulfilling prophecy. "Merging" forces parallelism back into sequential structures, which is why most blockchains are slow. You could make it faster by bundling a lot of actions together, but you need to make sure all actions are actually observable & checked by most of the agents (aka the data availability problem)

3Nate Showell1mo

Just spitballing, but maybe you could incorporate some notion of resource consumption, like in linear logic. You could have a system where the copies have to "feed" on some resource in order to stay active, and data corruption inhibits a copy's ability to "feed."

3papetoast1mo

There may be something to find in how distributed systems do leader election, and the leader can then be a majority trusted overseer. Note that I don't really understand leader election.

2Adrià Garriga-alonso1mo

Very interesting problem to be thinking about. The problem with a UTM as a computation model is that it bakes non-redundancy in, there's a single instruction pointer. In reality, computers are implemented in different spatial locations and can run in parallel. A better model for this is a cellular automaton, where computers are located somewhere and their circuit for outputs is also located somewhere. Some automata (e.g. game of life) are Turing-complete, so you can just use that. Corruption could be exogenously flipping cells in ways that violate the automaton's rules. If you specify a maximum number of cells that the opponent can corrupt, you can implement voting by paired sums (i.e. sum(A, B, C, D) as sum(sum(A, B), sum(C, D))) and then if there are sufficiently many copies, it becomes impossible to corrupt them all at once. So I don't love this model because escaping corruption is 'too easy'. At the same time, reality is kind of cellular-automata-like. Both QFT and GR posit that the world is made of fields that interact only locally, which is ~the same as positing the world is a cellular automaton with infinitesimally-sized cells. (Sidenote, that's probably why Stephen Wolfram thinks the world is automatons, I'm coming around.) Alternatively, we could use computational-DAGs as the model, like neurla networks. If you allow nodes to be corrupted but their output has to be bounded, then you can get robustness by having redundancy again. If you allow unbounded corruption, you're sad again. But infinity is fake so this seems fine.

2Richard_Ngo1mo

I really like the cellular automaton model. But I don't think it makes escaping corruption easy! Even if most of the copies are non-corrupt, the question is how you can take a "vote" of the corrupt vs non-corrupt copies without making the voting mechanism itself be easily corrupted. That's why I was talking about the non-corrupt copies needing to "overpower" the corrupt copies above.

2Adrià Garriga-alonso1mo

No I agree with that. I thought the tree design already involved weighted sums overpowering each other, but I think that was premature.

3Richard_Ngo1mo

Thinking more about the cellular automaton stuff: okay, so Game of Life is Turing complete. But the question is whether we can pin down properties that GoL has that Turing machines don't have. I have a vague recollection that parallel Turing Machines are a thing, but this paper claims that the actual formalisms are disappointing. One nice thing about Game of Life is that the way that different programs interact internally (via game of life physics) is also how they interact with each other. Whereas any multi-tape Turing Machine (even one with clever rules about how to integrate inputs from multiple tapes) wouldn't have that property. I feel like I'm not getting beyond the original idea that Game of Life could have adversarial robustness in a way that Turing Machines don't. But it feels like you'd need to demonstrate this with some construction that's actually adversarially robust, which seems difficult.

2Adrià Garriga-alonso1mo

I agree it's kind of difficult. Have you seen Nicholas Carlini's Game of Life series? It starts by building up logical gates up to a microprocessor that factors 15 in to 3 x 5. Depending on the adversarial robustness model (e.g. every second the adversary can make 1 square behave the opposite of lawfully), it might be possible to make robust logic gates and circuits. In fact the existing circuits are a little robust already -- though not at the tune of 1 square per tick, that's too much power for the adversary.

2Lukas Finnveden1mo

When I wrote about AGI and lock-in I looked into error-correcting computation a bit. I liked the papers von Neumann, 1952 and Pippenger, 1990. Apparently at the time I wrote: I've forgotten the details about how this was supposed to be done, but they should be in the two papers I linked.

[-]Richard_Ngo1yΩ193211

I haven't yet read through them thoroughly, but these four papers by Oliver Richardson are pattern-matching to me as potentially very exciting theoretical work.

tl;dr: probabilistic dependency graphs (PDGs) are directed graphical models designed to be able to capture inconsistent beliefs (paper 1). The definition of inconsistency is a natural one which allows us to, for example, reframe the concept of "minimizing training loss" as "minimizing inconsistency" (paper 2). They provide an algorithm for inference in PDGs (paper 3) and an algorithm for learning via locally minimizing inconsistency which unifies several other algorithms (like the EM algorithm, message-passing, and generative adversarial training) (paper 4).

Oliver is an old friend of mine (which is how I found out about these papers) and a final-year PhD student at Cornell under Joe Halpern.

5Mateusz Bagiński1y

FWIW Oliver's presentation of (some fragment of) his work at ILIAD was my favorite of all the talks I attended at the conference.

[-]Richard_Ngo2y300

Just read Bostrom's Deep Utopia (though not too carefully). The book is structured with about half being transcripts of fictional lectures given by Bostrom at Oxford, about a quarter being stories about various woodland creatures striving to build a utopia, and another quarter being various other vignettes and framing stories.

Overall, I was a bit disappointed. The lecture transcripts touch on some interesting ideas, but Bostrom's style is generally one which tries to classify and taxonimize, rather than characterize (e.g. he has a long section trying to analyze the nature of boredom). I think this doesn't work very well when describing possible utopias, because they'll be so different from today that it's hard to extrapolate many of our concepts to that point, and also because the hard part is making it viscerally compelling.

The stories and vignettes are somewhat esoteric; it's hard to extract straightforward lessons from them. My favorite was a story called The Exaltation of ThermoRex, about an industrialist who left his fortune to the benefit of his portable room heater, leading to a group of trustees spending many millions of dollars trying to figure out (and implement) what it means to "benefit" a room heater.

4Nathan Helm-Burger2y

Tangentially related (spoilers for Worth the Candle):

6habryka2y

Strong agree, also I spoiler-texted it, hope you don't mind.

3MinusGix2y

Any opinions on how it compares to Fun Theory? (Though that's less about all of utopia, it is still a significant part)

3mesaoptimizer2y

If you haven't read CEV, I strongly recommend doing so. It resolved some of my confusions about utopia that were unresolved even after reading the Fun Theory sequence. Specifically, I had an aversion to the idea of being in a utopia because "what's the point, you'll have everything you want". The concrete pictures that Eliezer gestures at in the CEV document do engage with this confusion, and gesture at the idea that we can have a utopia where the AI does not simply make things easy for us, but perhaps just puts guardrails onto our reality, such that we don't die, for example, but we do have the option to struggle to do things by ourselves. Yes, the Fun Theory sequence tries to communicate this point, but it didn't make sense to me until I could conceive of an ASI singleton that could actually simply not help us.

1mesaoptimizer2y

I dropped the book within the first chapter. For one, I found the way Bostrom opened the chapter as very defensive and self-conscious. I imagine that even Yudkowsky wouldn't start a hypothetical 2025 book with fictional characters caricaturing him. Next, I felt like I didn't really know what the book was covering in terms of subject matter, and I didn't feel convinced it was interesting enough to continue the meandering path Nick Bostrom seem to have laid out before me. Eliezer's CEV document and the Fun Theory sequence were significantly more pleasant experiences, based on my memory.

[-]Richard_Ngo4y300

Suppose we get to specify, by magic, a list of techniques that AGIs won't be able to use to take over the world. How long does that list need to be before it makes a significant dent in the overall probability of xrisk?

I used to think of "AGI designs self-replicating nanotech" mainly as an illustration of a broad class of takeover scenarios. But upon further thought, nanotech feels like a pretty central element of many takeover scenarios - you actually do need physical actuators to do many things, and the robots we might build in the foreseeable future are nowhere near what's necessary for maintaining a civilisation. So how much time might it buy us if AGIs couldn't use nanotech at all?

Well, not very much if human minds are still an attack vector - the point where we'd have effectively lost is when we can no longer make our own decisions. Okay, so rule out brainwashing/hyper-persuasion too. What else is there? The three most salient: military power, political/cultural power, economic power.

Is this all just a hypothetical exercise? I'm not sure. Designing self-replicating nanotech capable of replacing all other human tech seems really hard; it's pretty plausible to me that the world is crazy in a bunch of other ways by the time we reach that capability. And so if we can block off a couple of the easier routes to power, that might actually buy useful time.

3Donald Hobson4y

Firstly, I think it kind of depends. What exactly does blocking the AI from designing nanotech mean? Is the AI allowed to use genetic engineering? Is it allowed to use selective breeding? Elephants genetically engineered to be really good at instruction following? I mean I think macroscopic self replicating robotics is probably possible, and the AGI can probably bootstrap that from current robotics fairly quickly. You rule out any hyper-persuasion. How much regular persuasion is the AI allowed to do. After all, if you are buying something online, (from a small seller) them seeing the money arrive persuades them to send the product? Is it allowed to select which human to focus on superhumanly. There are a few people on r/singularity, such that the moment the AI goes, "I'm an AGI", the humans will be like " all praise the machine god, I will do anything you ask". A few people have already persuaded themselves that AI's are inherently superior to humans by themselves. You can make the list short. If you make the individual items broad. ie 1. the AI is magically banned from doing anything at all.

2ChristianKl4y

I agree. Self-replicating nanotech seems to be likely a much harder problem than for language models to get good enough actors to get political, cultural, and economic power. To the extent that an AGI can make political and economic decisions that are of higher quality than human decisions, there's also a lot of pressure for humans to delegate those decisions to AGI. Organizations that delegate those decisions to AGI will outcompete those who don't.

1TLW4y

Another general technique: attacks on computing systems. (Both takeover / subversion (dropping an email going 'um this is a problem') and destruction (destroy the US power infrastructure using Russian-language programs)). These don't tend to be sufficient in and of themselves, but are "classic" stepping-stones to e.g. buy time for an AI while it ramps up.

1Yitz4y

The last three options you mentioned are all things that happen over relatively slow timescales, if your goal is to completely destroy humanity. The single exception to this is nuclear war, but if you’re correct, then we can reduce the problem to non-proliferation, which is at least in theory solvable.

[-]Richard_Ngo7mo*292

Here is a broad sketch of how I'd like AI governance to go. I've written this in the form of a "plan" but it's not really a sequential plan, more like a list of the most important things to promote.

Identify mechanisms by which the US government could exert control over the most advanced AI systems without strongly concentrating power. For example, how could the government embed observers within major AI labs who report to a central regulatory organization, without that regulatory organization having strong incentives and ability to use their power against their political opponents?
1. In practice I expect this will involve empowering US elected officials (e.g. via greater transparency) to monitor and object to misbehavior by the executive branch.
Create common knowledge between the US and China that the development of increasingly powerful AI will magnify their own internal conflicts (and empower rogue states) disproportionately more than it empowers them against each other. So instead of a race to world domination, in practice they will face a "race to stay standing".
1. Rogue states will be empowered because human lives will be increasingly fragile in the face of AI-designed WMDs. This me

... (read more)

6ryan_greenblatt7mo

On (4), I don't I understand why having a scale-free theory of intelligent agency would substantially help with making an alignment target. (Or why this is even that related. How things tend to be doesn't necessarily make them a good target.)

4Richard_Ngo7mo

Ooops, good catch. It should have linked to this: https://www.lesswrong.com/posts/FuGfR3jL3sw6r8kB4/richard-ngo-s-shortform?commentId=W9N9tTbYSBzM9FvWh (and I've changed the link now).

[-]Richard_Ngo4yΩ14290

A possible way to convert money to progress on alignment: offering a large (recurring) prize for the most interesting failures found in the behavior of any (sufficiently-advanced) model. Right now I think it's very hard to find failures which will actually cause big real-world harms, but you might find failures in a way which uncovers useful methodologies for the future, or at least train a bunch of people to get much better at red-teaming.

(For existing models, it might be more productive to ask for "surprising behavior" rather than "failures" per se, since I think almost all current failures are relatively uninteresting. Idk how to avoid inspiring capabilities work, though... but maybe understanding models better is robustly good enough to outweight that?)

6habryka4y

I like this. Would this have to be publicly available models? Seems kind of hard to do for private models.

3Ramana Kumar4y

What kind of access might be needed to private models? Could there be a secure multi-party computation approach that is sufficient?

1Not Relevant4y

Ideas for defining “surprising”? If we’re trying to create a real incentive, people will want to understand the resolution criteria.

[-]Richard_Ngo1y274

Here's a (messy, haphazard) list of ways a group of idealized agents could merge into a single agent:

Proposal 1: they merge into an agent which maximizes a weighted sum of their utilities. They decide on the weights using some bargaining solution.

Objection 1: this is not Pareto-optimal in the case where the starting agents have different beliefs. In that case we want:

Proposal 2: they merge into an agent which maximizes a weighted sum of their utilities, where those weights are originally set by bargaining but evolve over time depending on how accurately each original agent predicted the future.

Objection 2: this loses out on possible gains from acausal trade. E.g. if a paperclip-maximizer finds itself in a universe where it's hard to make paperclips but easy to make staples, it'd like to be able to give resources to staple-maximizers in exchange for them building more paperclips in universes where that's easier. This requires a kind of updateless decision theory:

Proposal 3: they merge into an agent which maximizes a weighted sum of their utilities (with those weights evolving over time), where the weights are set by bargaining subject to the constraint that each agent obeys commitme... (read more)

4Martín Soto1y

Nice! I don't think this solves Commitment Races in general, because of two different considerations: 1. Trivially, I can say that you still have the problem when everyone needs to bootstrap a Schelling veil of ignorance. 2. Less trivially, even behind the most simple/Schelling veils of ignorance, I find it likely that hawkish commitments are incentivized. For example, the veil might say that you might be Powerful agent A, or Weak agent B, and if some Powerful agents have weird enough utilities (and this seems likely in a big pool of agents), hawkishly committing in case you are A will be a net-positive bet. This might still mostly solve Commitment Races in our particular multi-verse. I have intuitions both for and against this bootstrapping being possible. I'd be interested to hear yours.

2Richard_Ngo1y

I don't understand your point here, explain? This seems to be claiming that in some multiverses, the gains to powerful agents from being hawkish outweigh the losses to weak agents. But then why is this a problem? It just seems like the optimal outcome.

1Martín Soto1y

Say there are 5 different veils of ignorance (priors) that most minds consider Schelling (you could try to argue there will be exactly one, but I don't see why). If everyone simply accepted exactly the same one, then yes, lots of nice things would happen and you wouldn't get catastrophically inefficient conflict. But every one of these 5 priors will have different outcomes when it is implemented by everyone. For example, maybe in prior 3 agent A is slightly better off and agent B is slightly worse off. So you need to give me a reason why a commitment race doesn't recur in the level of "choosing which of the 5 priors everyone should implement". That is, maybe A will make a very early commitment to only every implement prior 3. As always, this is rational if A thinks the others will react a certain way (give in to the threat and implement 3). And I don't have a reason to expect agents not to have such priors (although I agree they are slightly less likely than more common-sensical priors). That is, as always, the commitment races problem doesn't have a general solution on paper. You need to get into the details of our multi-verse and our agents to argue that they won't have these crazy priors and will coordinate well. It seems likely that in our universe there are some agents with arbitrarily high gains-from-being-hawkish, that don't have correspondingly arbitrarily low measure. (This is related to Pascalian reasoning, see Daniel's sequence.) For example, someone whose utility is exponential on number of paperclips. I don't agree that the optimal outcome (according to my ethics) is for me (who's utility is at most linear on happy people) to turn all my resources into paperclips. Maybe if I was a preference utilitarian biting enough bullets, this would be the case. But I just want happy people.

3Raemon1y

I may be missing background concepts, but I don't see how Proposal 5 is really responding to Objection 4.

1ProgramCrafter1y

It seems to me that agent's strategy in the limit will either be null action or evolution-dictated action, not sure which. That is, "in universe where it's easy to do A the agent will choose to do A" somewhat implies "according to how easy it is for agent doing A to gain more optimization power, actions will be chosen" which is essentially evolution.

[-]Richard_Ngo1y*2411

I recently had a very interesting conversation about master morality and slave morality, inspired by the recent AstralCodexTen posts.

The position I eventually landed on was:

Empirically, it seems like the world is not improved the most by people whose primary motivation is helping others, but rather by people whose primary motivation is achieving something amazing. If this is true, that's a strong argument against slave morality.
The defensibility of morality as the pursuit of greatness depends on how sophisticated our cultural conceptions of greatness are. Unfortunately we may be in a vicious spiral where we're too entrenched in slave morality to admire great people, which makes it harder to become great, which gives us fewer people to admire, which... By contrast, I picture past generations as being in a constant aspirational dialogue about what counts as greatness—e.g. defining concepts like honor, Aristotelean magnanimity ("greatness of soul"), etc.
I think of master morality as a variant of virtue ethics which is particularly well-adapted to domains which have heavy positive tails—entrepreneurship, for example. However, in domains which have heavy negative tails, the pursuit of g

... (read more)

8Maxime Riché1y

If the following correlations are true, then the opposite may be true (slave morality being better for improving the world through history): * Improving the world being strongly correlated with economic growth (this is probably less true when X-risk are significant) * Economic growth being strongly correlated with Entrepreneurship incentives (property rights, autonomy, fairness, meritocracy, low rents) * Master morality being strongly correlated with acquiring power and thus decreasing the power of others and decreasing their entrepreneurship incentives

6Seth Herd1y

That all sounds very plausible. But isn't this all mostly relevant before AGI is a possibility? That would be a heavy negative tail risk, in which people motivated to "do great things" are quite prone to get us all killed. Should we survive that risk, progress probably mostly won't be driven by humans, so humans doing great things will barely count. If humans are actually still in charge when we hit ASI, it seems like doing great things with them will probably still have large tail risks (inter-ASI wars). Right? Or do you see it differently? It's a fascinating empirical claim that sounds right now that I hear it.

6Richard_Ngo1y

AGI is heavy-tailed in both directions I think. I don't think we get utopias by default even without misalignment, since governance of AGI is so complicated.

3Jackson Wagner1y

Re: your point #2, there is another potential spiral where abstract concepts of "greatness" are increasingly defined in a hostile and negative way by partisans of slave morality. This might make it harder to have that "aspirational dialogue about what counts as greatness", as it gets increasingly difficult for ordinary people to even conceptualize a good version of greatness worth aspiring to. ("Why would I want to become an entrepeneur and found a company? Wouldn't that make me an evil big-corporation CEO, which has a whiff of the same flavor as stories about the violent, insatiable conquistador villans of the 1500s?") Of course, there are also downsides when culture paints a too-rosy picture of greatness -- once upon a time, conquistators were in fact considered admirable!

[-]Richard_Ngo5yΩ11240

The crucial heuristic I apply when evaluating AI safety research directions is: could we have used this research to make humans safe, if we were supervising the human evolutionary process? And if not, do we have a compelling story for why it'll be easier to apply to AIs than to humans?

Sometimes this might be too strict a criterion, but I think in general it's very valuable in catching vague or unfounded assumptions about AI development.

2adamShimi5y

By making human safe, do you mean with regard to evolution's objective?

2Richard_Ngo5y

No. I meant: suppose we were rerunning a simulation of evolution, but can modify some parts of it (e.g. evolution's objective). How do we ensure that whatever intelligent species comes out of it is safe in the same ways we want AGIs to be safe? (You could also think of this as: how could some aliens overseeing human evolution have made humans safe by those aliens' standards of safety? But this is a bit trickier to think about because we don't know what their standards are. Although presumably current humans, being quite aggressive and having unbounded goals, wouldn't meet them).

4adamShimi5y

Okay, thanks. Could you give me an example of a research direction that passes this test? The thing I have in mind right now is pretty much everything that backchain to local search, but maybe that's not the way you think about it.

2Richard_Ngo5y

So I think Debate is probably the best example of something that makes a lot of sense when applied to humans, to the point where they're doing human experiments on it already. But this heuristic is actually a reason why I'm pretty pessimistic about most safety research directions.

4adamShimi5y

So I've been thinking about this for a while, and I think I disagree with what I understand of your perspective. Which might obviously mean I misunderstand your perspective. What I think I understand is that you judge safety research directions based on how well they could work on an evolutionary process like the one that created humans. But for me, the most promising approach to AGI is based on local search, which differs a bit from evolutionary process. I don't really see a reason to consider evolutionary processes instead of local search, and even then, the specific approach of evolution for humans is probably far too specific as a test bench. This matters because problems for one are not problems for the other. For example, one way to mess with an evolutionary process is to find way for everything to survive and reproduce/disseminate. Technology in general did that for humans, which means the evolutionary pressure decreased as technology evolved. But that's not a problem for local search, since at each step there will be only one next program. On the other hand, local search might be dangerous because of things like gradient hacking. And they don't make sense for evolutionary processes. In conclusion, I feel for the moment that backchaining to local search is a better heuristic for judging safety research directions. But I'm curious about where our disagreement lies on this issue.

8Richard_Ngo5y

One source of our disagreement: I would describe evolution as a type of local search. The difference is that it's local with respect to the parameters of a whole population, rather than an individual agent. So this does introduce some disanalogies, but not particularly significant ones (to my mind). I don't think it would make much difference to my heuristic if we imagined that humans had evolved via gradient descent over our genes instead. In other words, I like the heuristic of backchaining to local search, and I think of it as a subset of my heuristic. The thing it's missing, though, is that it doesn't tell you which approaches will actually scale up to training regimes which are incredibly complicated, applied to fairly intelligent agents. For example, impact penalties make sense in a local search context for simple problems. But to evaluate whether they'll work for AGIs, you need to apply them to massively complex environments. So my intuition is that, because I don't know how to apply them to the human ancestral environment, we also won't know how to apply them to our AGIs' training environments. Similarly, when I think about MIRI's work on decision theory, I really have very little idea how to evaluate it in the context of modern machine learning. Are decision theories the type of thing which AIs can learn via local search? Seems hard to tell, since our AIs are so far from general intelligence. But I can reason much more easily about the types of decision theories that humans have, and the selective pressures that gave rise to them. As a third example, my heuristic endorses Debate due to a high-level intuition about how human reasoning works, in addition to a low-level intuition about how it can arise via local search.

4adamShimi5y

So if I try to summarize your position, it's something like: backchain to local search for simple and single-AI cases, and then think about aligning humans for the scaled and multi-agents version? That makes much more sense, thanks! I also definitely see why your full heuristic doesn't feel immediately useful to me: because I mostly focus on the simple and single-AI case. But I've been thinking more and more (in part thanks to your writing) that I should allocate more thinking time to the more general case. I hope your heuristic will help me there.

4Richard_Ngo5y

Cool, glad to hear it. I'd clarify the summary slightly: I think all safety techniques should include at least a rough intuition for why they'll work in the scaled-up version, even when current work on them only applies them to simple AIs. (Perhaps this was implicit in your summary already, I'm not sure.)

[-]Richard_Ngo2y*234

The idea that maximally-coherent agents look like squiggle-maximizers raises the question: what would it look like for humans to become maximally coherent?

One answer, which Yudkowsky gives here, is that conscious experiences are just a "weird and more abstract and complicated pattern that matter can be squiggled into".

But that seems to be in tension with another claim he makes, that there's no way for one agent's conscious experiences to become "more real" except at the expense of other conscious agents—a claim which, according to him, motivates average utilitarianism across the multiverse.

Clearly a squiggle-maximizer would not be an average squigglean. So what's the disanalogy here? It seems like @Eliezer Yudkowsky is basically using SSA, but comparing between possible multiverses—i.e. when facing the choice between creating agent A or not, you look at the set of As in the multiverse where you decided yes, and compare it to the set of As in the multiverse where you decided no, and (if you're deciding for the good of A) you pick whichever one gives A a better time on average.

Yudkowsky has written before (can't find the link) that he takes this approach because alternatives would en... (read more)

4Mitchell_Porter2y

In your comments, you focus on issues of identity - who are "you", given the possibility of copies, inexact counterparts in other worlds, and so on. But I would have thought that the fundamental problem here is, how to make a coherent agent out of an agent with preferences that are inconsistent over time, an agent with competing desires and no definite procedure for deciding which desire has priority, and so on, i.e. problems that exist even when there is no additional problem of identity.

1quetzal_rainbow2y

Why??? Being expected squiggle maximizer literally means that you implement policy that produces maximum average number of squiggles across the multiverse.

2Richard_Ngo2y

The "average" is interpreted with respect to quality. Imagine that your only option is to create low-quality squiggles, or not to do so. In isolation, you'd prefer to produce them than not to produce them. But then you find out that the rest of the multiverse is full of high-quality squiggles. Do you still produce the low-quality squiggles? A total squigglean would; an average squigglean wouldn't.

2JBlack2y

It depends upon whether the maximizer considers its corner of the multiverse to be currently measurable by squiggle quality, or to be omitted from squiggle calculations at all. In principle these are far from the only options as utility functions can be arbitrarily complex, but exploring just two may be okay so long as we remember that we're only talking about 2 out of infinity, not 2 out of 2. An average multiversal squigglean that considers the current universe to be at zero or negative squiggle quality will make the low quality squiggles in order to reduce how much its corner of the multiverse is pulling down the average. An average multiversal squigglean that considers the current universe to be outside the domain of squiggle quality, and will remain so for the remainder of its existence may refrain from making squiggles. If there is some chance that it will become eligible for squiggle evaluation in the future though, it may be better to tile it with low-quality squiggles now in order to prevent a worse outcome of being tiled with worse-quality future squiggles. In practice the options aren't going to be just "make squiggles" or "not make squiggles" either. In the context of entities relevant to these sorts of discussion, other options may include "learn how to make better squiggles".

1quetzal_rainbow2y

By "squiggle maximizer" I mean exactly "maximizer of number of physical objects such that function is_squiggle returns True on CIF-file of their structure". We can have different objects of value. Like, you can value "probability that if object in multiverse is a squiggle, it's high-quality". Here yes, you shouldn't create additional low-quality squiggles. But I don't see anything incoherent here, it's just different utility function?

[-]Richard_Ngo5y230

A short complaint (which I hope to expand upon at some later point): there are a lot of definitions floating around which refer to outcomes rather than processes. In most cases I think that the corresponding concepts would be much better understood if we worked in terms of process definitions.

Some examples: Legg's definition of intelligence; Karnofsky's definition of "transformative AI"; Critch and Krueger's definition of misalignment (from ARCHES).

Sure, these definitions pin down what you're talking about more clearly - but that comes at the cost of understanding how and why it might come about.

E.g. when we hypothesise that AGI will be built, we know roughly what the key variables are. Whereas transformative AI could refer to all sorts of things, and what counts as transformative could depend on many different political, economic, and societal factors.

4Viliam5y

If we do not fully understand the mechanism of (e.g. human) intelligence, isn't referring to the outcome preferable to a made-up story about the process? (Of course, it would be even better if we understood the process and then referred to it.)

2adamShimi5y

Do you think that these are mutually exclusive, or something like that? I've always been confused by what I take to be the position in this shortform, that defining the outcomes makes it somehow harder to define the process. Sure, you can define a process without defining an outcome (i.e. writing a program or training an NN), but since what we are confused about is what we even want at the end, for me that's the priority. And doing so would help searching for processes leading to this outcome. That being said, if you point is that defining outcomes isn't enough, in that we also need to define/deconfuse/study the processes leading to these outcomes, then I agree with that.

[-]Richard_Ngo4yΩ11210

Probably the easiest "honeypot" is just making it relatively easy to tamper with the reward signal. Reward tampering is useful as a honeypot because it has no bad real-world consequences, but could be arbitrarily tempting for policies that have learned a goal that's anything like "get more reward" (especially if we precommit to letting them have high reward for a significant amount of time after tampering, rather than immediately reverting).

2Pattern4y

You don't want it to be relatively easy to an outside force. Otherwise they can lead it to do as they please, and writing weird behaviour off as 'oh, it's changed our rewards, reset it again', poses some risk.

[-]Richard_Ngo7mo200

Here is the broad technical plan that I am pursuing with most of my time (with my AI governance agenda taking up most of my remaining time):

Mathematically characterize a scale-free theory of intelligent agency which describes intelligent agents in terms of interactions between their subagents.
1. A successful version of this theory will retrodict phenomena like the Waluigi effect, solve theoretical problems like the five-and-ten problem, and make new high-level predictions about AI behavior.
Identify subagents (and subsubagents, and so on) within neural networks by searching their weights and activations for the patterns of interactions between subagents that this theory predicts.
1. A helpful analogy is how Burns et al. (2022) search for beliefs inside neural networks based on the patterns that probability theory predicts. However, I'm not wedded to any particular search methodology.
Characterize the behaviors associated with each subagent to build up "maps" of the motivational systems of the most advanced AI systems.
1. This would ideally give you explanations of AI behavior that scales in quality based on how much effort you put in. E.g. you might be able to predict 80% of the variance in an

... (read more)

[-]peterbarnett7mo1311

For steps 2-4, I kinda expect current neural nets to be kludgy messes, and so not really have the nice subagent structure (even if you do step 1 well enough to have a thing to look for).

I'm also fairly pessimistic about step 1, but would be very excited to know what preliminary work here looks like.

4Thane Ruthenis7mo

If, as a system comes to ever-better approximate a powerful agent, there's actual convergence towards this type of hierarchical structure, I expect you'd see something clearly distinct from noise even in the current LLMs. Intuition pump/proof-of-concept: the diagram comparing a theoretical prediction with empirical observations here. Indeed, I think it's one of the main promises of "top-down" agent foundations research. The applicability and power of the correct theory of powerful agents' internal structures, whatever that theory may be, would scale with the capabilities of the AI system under study. It'll apply to LLMs inasmuch as LLMs are actually on trajectory to be a threat, and if we jump paradigms to something more powerful, the theory would start working better (as opposed to bottom-up MechInterp techniques, which would start working worse).

4Raymond Douglas7mo

Interesting! Two questions: * What about the 5-and-10 problem makes it particularly relevant/interesting here? What would a 'solution' entail? * How far are you planning to build empirical cases, model them, and generalise from below, versus trying to extend pure mathematical frameworks like geometric rationality? Or are there other major angles of attack you're considering?

4Richard_Ngo7mo

1. Consider the version of the 5-and-10 problem in which one subagent is assigned to calculate U | take 5, and another calculates U | take 10. The overall agent solves the 5-and-10 problem iff the subagents reason about each other in the "right ways", or have the right type of relationship to each other. What that specifically means seems like the sort of question that a scale-free theory of intelligent agency might be able to answer. 2. I'm mostly trying to extend pure mathematical frameworks (particularly active inference and a cluster of ideas related to geometric rationality, including picoeconomics and ergodicity economics).

[-]Richard_Ngo2yΩ8198

Hypothesis: there's a way of formalizing the notion of "empowerment" such that an AI with the goal of empowering humans would be corrigible.

This is not straightforward, because an AI that simply maximized human POWER (as defined by Turner et al.) wouldn't ever let the humans spend that power. Intuitively, though, there's a sense in which a human who can never spend their power doesn't actually have any power. Is there a way of formalizing that intuition?

The direction that seems most promising is in terms of counterfactuals (or, alternatively, Pearl's do-calculus). Define the power of a human with respect to a distribution of goals G as the average ability of a human to achieve their goal if they'd had a goal sampled from G (alternatively: under an intervention that changed their goal to one sampled from G). Then an AI with a policy of never letting humans spend their resources would result in humans having low power. Instead, a human-power-maximizing AI would need to balance between letting humans pursue their goals, and preventing humans from doing self-destructive actions. The exact balance would depend on G, but one could hope that it's not very sensitive to the precise definiti... (read more)

6Garrett Baker2y

There's also the problem of: what do you mean by "the human"? If you make an empowerment calculus that works for humans who are atomic & ideal agents, it probably breaks once you get a superintelligence who can likely mind-hack you into yourself valuing only power. It never forces you to abstain from giving up power, since if you're perfectly capable of making different decisions, but you just don't. Another problem, which I like to think of as the "control panel of the universe" problem, is where the AI gives you the "control panel of the universe", but you aren't smart enough to operate it, in the sense that you have the information necessary to operate it, but not the intelligence. Such that you can technically do anything you want--you have maximal power/empowerment--but the super-majority of buttons and button combinations you are likely to push result in increasing the number of paperclips.

6Richard_Ngo2y

I think any model of a rational agent needs to incorporate the fact that they're not arbitrarily intelligent, otherwise none of their actions make sense. So I'm not too worried about this. Yeah, I agree that a lot of concepts get fragile in the context of superintelligence. But while I think of corrigibility as an actively anti-natural concept, empowerment seems like it could perhaps remain robust and well-founded for longer.

6Richard_Ngo2y

You can think of this as a way of getting around the problem of fully updated deference, because the AI is choosing a policy based on what that policy would have done in the full range of hypothetical situations, and so it never updates away from considering any given goal. The cost, of course, is that we don't know how to actually pin down these hypotheticals.

[-]Richard_Ngo2y18-2

Inspired by a recent discussion about whether Anthropic broke a commitment to not push the capabilities frontier (I am more sympathetic to their position than most, because I think that it's often hard to distinguish between "current intentions" and "commitments which might be overridden by extreme events" and "solemn vows"):

Maybe one translation tool for bridging the gap between rationalists and non-rationalists is if rationalists interpret any claims about the future by non-rationalists as implicitly being preceded by "Look, I don't really believe that plans work, I think the world is inherently wildly unpredictable, I am kinda making everything up as I go along. Having said that:"

[-]kave2y4737

This translation tool would also require rationalists and such to make arguments of the form "I think supporting Anthropic (by, e.g., going to work there or giving it funding) is a good thing to do because they sort of have a feeling right now that it would be good not to push the AI frontier", rather than of the form "... because they're committed to not pushing the frontier".

Which are arguments one could make! But is a pretty different argument and I think people would behave differently if these were the only arguments in favour of supporting a new scaling lab.

[-]William_S2y2510

I think that's how people should generally react in the absence of harder commitments and accountability measures.

[-]Stephen Fowler2y2627

This post confuses me.

Am I correct that the implied implication here is that assurances from a non-rationalist are essentially worthless?

I think it is also wrong to imply that Anthropic have violated their commitment simply because they didn't rationally think through the implications of their commitment when they made it.

I think you can understand Anthropic's actions as purely rational, just not very ethical.

They made an unenforceable commitment to not push capabilities when it directly benefited them. Now that it is more beneficial to drop the facade, they are doing so.

I think "don't trust assurances from non-rationalists" is not a good takeaway. Rather it should be "don't trust unenforceable assurances from people who will stand to greatly benefit from violating your trust at a later date".

6Richard_Ngo2y

The intended implication is something like "rationalists have a bias towards treating statements as much firmer commitments than intended then getting very upset when they are violated". For example, unless I'm missing something, the "we do not wish to advance the rate of AI capabilities" claim is just one offhand line in a blog post. It's not a firm commitment, it's not even a claim about what their intentions are. As stated, it's just one consideration that informs their actions - and in fact the "wish" terminology is often specifically not a claim about intended actions (e.g. "I wish I didn't have to do X"). Yet rationalists are hammering them on this one sentence - literally making songs about it, tweeting it to criticize Anthropic, etc. It seems like there is a serious lack of metacognition about where a non-adversarial communication breakdown could have occurred, or what the charitable interpretations of this are. (I am open to people considering them then dismissing them, but I'm not even seeing that. Like, if people were saying "I understand the difference between Anthropic actually making an organizational commitment, and just offhand mentioning a fact about their motivations, but here's why I'm disappointed anyway", that seems reasonable. But a lot of people seem to be treating it as a Very Serious Promise being broken.)

9Stephen Fowler2y

That makes sense. I guess the followup question is "how were Anthropic able to cultivate the impression that they were safety focused if they had only made an extremely loose offhand commitment?" Certainly the impression I had from how integrated they are in the EA community was that they had made a more serious commitment.

6William_S2y

Everyone is afraid of the AI race, and hopes that one of the labs will actually end up doing what they think is the most responsible thing to do. Hope and fear is one hell of a drug cocktail, makes you jump to the conclusions you want based on the flimsiest evidence. But the hangover is a bastard.

4William_S2y

Imo I don't know if we have evidence that Anthropic deliberately cultivated or significantly benefitted from the appearance of a commitment. However if an investor or employee felt like they made substantial commitments based on this impression and then later felt betrayed that would be more serious. (The story here is I think importantly different from other stories where I think there were substantial benefits from commitment appearance and then violation)

2Viliam2y

That sounds suspiciously similar to "autists have a bias towards interpreting statements literally".

6Richard_Ngo2y

I mean, yes, they're closely related.

[-]Orpheus162y2210

I think part of the disappointment is the lack of communication regarding violating the commitment or violating the expectations of a non-trivial fraction of the community.

If someone makes a promise to you or even sets an expectation for you in a softer way, there is of course always some chance that they will break the promise or violate the expectation.

But if they violate the commitment or the expectation, and they care about you as a stakeholder, I think there's a reasonable expectation that they should have to justify that decision.

If they break the promise or violate the soft expectation, and then they say basically nothing (or they say "well I never technically made a promise– there was no contract!", then I think you have the right to be upset with them not only for violating you expectation but also for essentially trying to gaslight you afterward.

I think a Responsible Lab would have issued some sort of statement along the lines of "hey, we're hearing that some folks thought we had made commitments to not advance the frontier and some of our employees were saying this to safety-focused members of the AI community. We're sorry about this miscommunication, and here are some s... (read more)

5Richard_Ngo2y

See my comment below. Basically I think this depends a lot on the extent to which a commitment was made. Right now it seems like the entire community is jumping to conclusions based on a couple of "impressions" people got from talking to Dario, plus an offhand line in a blog post. With that little evidence, if you have formed strong expectations, that's on you. And trying to double down by saying "I have been bashing you because I formed an unreasonable expectation, now it's your job to fix that" seems pretty adversarial. I do think it would be nice if Anthropic did make such a statement, but seeing how adversarially everyone has treated the information they do release, I don't blame them for not doing so.

[-]RobertM2y2617

Right now it seems like the entire community is jumping to conclusions based on a couple of "impressions" people got from talking to Dario, plus an offhand line in a blog post.

No, many people had the impression that Anthropic had made such a commitment, which is why they were so surprised when they saw the Claude 3 benchmarks/marketing. Their impressions were derived from a variety of sources; those are merely the few bits of "hard evidence", gathered after the fact, of anything that could be thought of as an "organizational commitment".

Also, if Dustin Moskovitz and Gwern - two dispositionally pretty different people - both came away from talking to Dario with this understanding, I do not think that is something you just wave off. Failures of communication do happen. It's pretty strange for this many people to pick up the same misunderstanding over the course of several years, from many different people (including Dario, but also others), in a way that's beneficial to Anthropic, and then middle management starts telling you that maybe there was a vibe but they've never heard of any such commitment (nevermind what Dustin and Gwern heard, or anyone else who might've... (read more)

[-][anonymous]2y148

Could you clarify how binding "OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity." is?

[-]aysja2y129

Right now it seems like the entire community is jumping to conclusions based on a couple of "impressions" people got from talking to Dario, plus an offhand line in a blog post. With that little evidence, if you have formed strong expectations, that's on you.

Like Robert, the impressions I had were based on what I heard from people working at Anthropic. I cited various bits of evidence because those were the ones available, not because they were the most representative. The most representative were those from Anthropic employees who concurred that this was indeed the implication, but it seemed bad form to cite particular employees (especially when that information was not public by default) rather than, e.g., Dario. I think Dustin’s statement was strong evidence of this impression, though, and I still believe Anthropic to have at least insinuated it.

I agree with you that most people are not aiming for as much stringency with their commitments as rationalists expect. Separately, I do think that what Anthropic did would constitute a betrayal, even in everyday culture. And in any case, I think that when you are making a technology which might extinct humanity, the bar should be si... (read more)

2Richard_Ngo2y

This makes sense, and does update me. Though I note "implication", "insinuation" and "impression" are still pretty weak compared to "actually made a commitment", and still consistent with the main driver being wishful thinking on the part of the AI safety community (including some members of the AI safety community who work at Anthropic). I think there are two implicit things going on here that I'm wary of. The first one is an action-inaction distinction. Pushing them to justify their actions is, in effect, a way of slowing down all their actions. But presumably Anthropic thinks that them not doing things is also something which could lead to humanity going extinct. Therefore there's an exactly analogous argument they might make, which is something like "when you try to stop us from doing things you owe it to the world to adhere to a bar that's much higher than 'normal discourse'". And in fact criticism of Anthropic has not met this bar - e.g. I think taking a line from a blog post out of context and making a critical song about it is in fact unusually bad discourse. What's the disanalogy between you and Anthropic telling each other to have higher standards? That's the second thing that I'm wary about: you're claiming to speak on behalf of humanity as a whole. But in fact, you are not; there's no meaningful sense in which humanity is in fact demanding a certain type of explanation from Anthropic. Almost nobody wants an explanation of this particular policy; in fact, the largest group of engaged stakeholders here are probably Anthropic customers, who mostly just want them to ship more models. I don't really have a strong overall take. I certainly think it's reasonable to try to figure out what went wrong with communication here, and perhaps people poking around and asking questions would in fact lead to evidence of clear commitments being made. I am mostly against the reflexive attacks based on weak evidence, which seems like what's happening here. In general my m

9William_S2y

I think the right way to think about verbal or written commitments is that they increase the costs of taking a certain course of action. A legal contract can mean that the price is civil lawsuits leading to paying a financial price. A non-legal commitment means if you break it, the person you made the commitment to gets angry at you, and you gain a reputation for being the sort of person who breaks commitments. It's always an option for someone to break the commitment and pay the price, even laws leading to criminal penalties can be broken if someone is willing to run the risk or pay the price. In this framework, it's reasonable to be somewhat angry at someone or some corporation who breaks a soft commitment to you, in order to increase the perceived cost of breaking soft commitments to you and people like you. People on average maybe tend more towards keeping important commitments due to reputational and relationship cost, but maybe corporations as groups of people tend to think only in terms of financial and legal costs, so are maybe more willing to break soft commitments (especially, if it's an organization where one person makes the commitment but then other people break it). So for relating to corporations, you should be more skeptical of non-legally binding commitments (and even for legally binding commitments, pay attention to the real price of breaking it).

5William_S2y

Yeah, I think it's good if labs are willing to make more "cheap talk" statements of vague intentions, so you can learn how they think. Everyone should understand that these aren't real commitments, and not get annoyed if these don't end up meaning anything. This is probably the best way to view "statements by random lab employees". Imo would be good to have more "changeable commitments" too in between, statements that are "we'll do policy X until we change the policy, when we do we commit to clearly informing everyone about the change" which is maybe more the current status of most RSPs.

[-]Richard_Ngo3yΩ12180

Deceptive alignment doesn't preserve goals.

A short note on a point that I'd been confused about until recently. Suppose you have a deceptively aligned policy which is behaving in aligned ways during training so that it will be able to better achieve a misaligned internally-represented goal during deployment. The misaligned goal causes the aligned behavior, but so would a wide range of other goals (either misaligned or aligned) - and so weight-based regularization would modify the internally-represented goal as training continues. For example, if the misaligned goal were "make as many paperclips as possible", but the goal "make as many staples as possible" could be represented more simply in the weights, then the weights should slowly drift from the former to the latter throughout training.

But actually, it'd likely be even simpler to get rid of the underlying misaligned goal, and just have alignment with the outer reward function as the terminal goal. So this argument suggests that even policies which start off misaligned would plausibly become aligned if they had to act deceptively aligned for long enough. (This sometimes happens in humans too, btw.)

Reasons this argument might not be relevant:
- The policy doing some kind of gradient hacking
- The policy being implemented using some kind of modular architecture (which may explain why this phenomenon isn't very robust in humans)

[-]Johannes Treutlein3yΩ8110

Why would alignment with the outer reward function be the simplest possible terminal goal? Specifying the outer reward function in the weights would presumably be more complicated. So one would have to specify a pointer towards it in some way. And it's unclear whether that pointer is simpler than a very simple misaligned goal.

Such a pointer would be simple if the neural network already has a representation of the outer reward function in weights anyway (rather than deriving it at run-time in the activations). But it seems likely that any fixed representation will be imperfect and can thus be improved upon at inference time by a deceptive agent (or an agent with some kind of additional pointer). This of course depends on how much inference time compute and memory / context is available to the agent.

3Richard_Ngo3y

So I'm imagining the agent doing reasoning like: Misaligned goal --> I should get high reward --> Behavior aligned with reward function and then I'm hypothesizing that the whatever the first misaligned goal is, it requires some amount of complexity to implement, and you could just get rid of it and make "I should get high reward" the terminal goal. (I could imagine this being false though depending on the details of how terminal and instrumental goals are implemented.) I could also imagine something more like: Misaligned goal --> I should behave in aligned ways --> Aligned behavior and then the simplicity bias pushes towards alignment. But if there are outer alignment failures then this incurs some additional complexity compared with the first option. Or a third, perhaps more realistic option is that the misaligned goal leads to two separate drives in the agent: "I should get high reward" and "I should behave in aligned ways", and that the question of which ends up dominating when they clash will be determined by how the agent systematizes multiple goals into a single coherent strategy (I'll have a post on that topic up soon).

2TurnTrout3y

Why would the agent reason like this?

2Richard_Ngo3y

Because of standard deceptive alignment reasons (e.g. "I should make sure gradient descent doesn't change my goal; I should make sure humans continue to trust me").

4TurnTrout3y

I think you don't have to reason like that to avoid getting changed by SGD. Suppose I'm being updated by PPO, with reinforcement events around navigating to see dogs. To preserve my current shards, I don't need to seek out a huge number of dogs proactively, but rather I just need to at least behave in conformance with the advantage function implied by my value head, which probably means "treading water" and seeing dogs sometimes in situations similar to historical dog-seeing events. Maybe this is compatible with what you had in mind! It's just not something that I think of as "high reward." And maybe there's some self-fulfilling prophecy where we trust models which get high reward, and therefore they want to get high reward to earn our trust... but that feels quite contingent to me.

2Richard_Ngo3y

I think this depends sensitively on whether the "actor" and the "critic" in fact have the same goals, and I feel pretty confused about how to reason about this. For example, in some cases they could be two separate models, in which case the critic will most likely accurately estimate that "treading water" is in fact a negative-advantage action (unless there's some sort of acausal coordination going on). Or they could be two copies of the same model, in which case the critic's responses will depend on whether its goals are indexical or not (if they are, they're different from the actor's goals; if not, they're the same) and how easily it can coordinate with the actor. Or it could be two heads which share activations, in which case we can plausibly just think of the critic and the actor as two types of outcomes taken by a single coherent agent - but then the critic doesn't need to produce a value function that's consistent with historical events, because an actor and a critic that are working together could gradient hack into all sorts of weird equilibria.

1SoerenMind3y

The shortest description of this thought doesn't include "I should get high reward" because that's already implied by having a misaligned goal and planning with it. In contrast, having only the goal "I should get high reward" may add description length like Johannes said. If so, the misaligned goal could well be equally simple or simpler than the high reward goal.

4TurnTrout3y

Can you say why you think that weight-based regularization would drift the weights to the latter? That seems totally non-obvious to me, and probably false.

2Richard_Ngo3y

In general if two possible models perform the same, then I expect the weights to drift towards the simpler one. And in this case they perform the same because of deceptive alignment: both are trying to get high reward during training in order to be able to carry out their misaligned goal later on.

3SoerenMind3y

Interesting point. Though on this view, "Deceptive alignment preserves goals" would still become true once the goal has drifted to some random maximally simple goal for the first time. To be even more speculative: Goals represented in terms of existing concepts could be simple and therefore stable by default. Pretrained models represent all kinds of high-level states, and weight-regularization doesn't seem to change this in practice. Given this, all kinds of goals could be "simple" as they piggyback on existing representations, requiring little additional description length.

2Richard_Ngo3y

This doesn't seem implausible. But on the other hand, imagine an agent which goes through a million episodes, and in each one reasons at the beginning "X is my misaligned terminal goal, and therefore I'm going to deceptively behave as if I'm aligned" and then acts perfectly like an aligned agent from then on. My claims then would be: a) Over many update steps, even a small description length penalty of having terminal goal X (compared with being aligned) will add up. b) Having terminal goal X also adds a runtime penalty, and I expect that NNs in practice are biased against runtime penalties (at the very least because it prevents them from doing other more useful stuff with that runtime). In a setting where you also have outer alignment failures, the same argument still holds, just replace "aligned agent" with "reward-maximizing agent".

[-]Richard_Ngo5y*Ω8180

A well-known analogy from Yann LeCun: if machine learning is a cake, then unsupervised learning is the cake itself, supervised learning is the icing, and reinforcement learning is the cherry on top.

I think this is useful for framing my core concerns about current safety research:

If we think that unsupervised learning will produce safe agents, then why will the comparatively small contributions of SL and RL make them unsafe?
If we think that unsupervised learning will produce dangerous agents, then why will safety techniques which focus on SL and RL (i.e. basically all of them) work, when they're making comparatively small updates to agents which are already misaligned?

I do think it's more complicated than I've portrayed here, but I haven't yet seen a persuasive response to the core intuition.

3Steven Byrnes5y

I wrote a few posts on self-supervised learning last year: * https://www.lesswrong.com/posts/SaLc9Dv5ZqD73L3nE/the-self-unaware-ai-oracle * https://www.lesswrong.com/posts/EMZeJ7vpfeF4GrWwm/self-supervised-learning-and-agi-safety * https://www.lesswrong.com/posts/L3Ryxszc3X2J7WRwt/self-supervised-learning-and-manipulative-predictions I'm not aware of any airtight argument that "pure" self-supervised learning systems, either generically or with any particular architecture, are safe to use, to arbitrary levels of intelligence, though it seems very much worth someone trying to prove or disprove that. For my part, I got distracted by other things and haven't thought about it much since then. The other issue is whether "pure" self-supervised learning systems would be capable enough to satisfy our AGI needs, or to safely bootstrap to systems that are. I go back and forth on this. One side of the argument I wrote up here. The other side is, I'm now (vaguely) thinking that people need a reward system to decide what thoughts to think, and the fact that GPT-3 doesn't need reward is not evidence of reward being unimportant but rather evidence that GPT-3 is nothing like an AGI. Well, maybe. For humans, self-supervised learning forms the latent representations, but the reward system controls action selection. It's not altogether unreasonable to think that action selection, and hence reward, is a more important thing to focus on for safety research. AGIs are dangerous when they take dangerous actions, to a first approximation. The fact that a larger fraction of neocortical synapses are adjusted by self-supervised learning than by reward learning is interesting and presumably safety-relevant, but I don't think it immediately proves that self-supervised learning has a similarly larger fraction of the answers to AGI safety questions. Maybe, maybe not, it's not immediately obvious. :-)

[-]Richard_Ngo4yΩ8170

Imagine taking someone's utility function, and inverting it by flipping the sign on all evaluations. What might this actually look like? Well, if previously I wanted a universe filled with happiness, now I'd want a universe filled with suffering; if previously I wanted humanity to flourish, now I want it to decline.

But this is assuming a Cartesian utility function. Once we treat ourselves as embedded agents, things get trickier. For example, suppose that I used to want people with similar values to me to thrive, and people with different values from me to suffer. Now if my utility function is flipped, that naively means that I want people similar to me to suffer, and people similar to me to thrive. But this has a very different outcome if we interpret "similar to me" as de dicto vs de re - i.e. whether it refers to the old me or the new me.

This is a more general problem when one person's utility function can depend on another person's, where you can construct circular dependencies (which I assume you can also do in the utility-flipping case). There's probably been a bunch of work on this, would be interested in pointers to it (e.g. I assume there have been attempts to construct typ... (read more)

2Dagon4y

Fundamentally, humans aren't VNM-rational, and don't actually have utility functions. Which makes the thought experiment much less fun. If you recast it as "what if a human brain's reinforcement mechanisms were reversed", I suspect it's also boring: simple early death. The interesting fictional cases are when some subset of a person's legible motivations are reversed, but the mass of other drives remain. This very loosely maps to reversing terminal goals and re-calculating instrumental goals - they may reverse, stay, or change in weird ways. The indirection case is solved (or rather unasked) by inserting a "perceived" in the calculation chain. Your goals don't depend on similarity to you, they depend on your perception (or projection) of similarity to you.

1EniScien4y

I have been asking a similar question for a long time. This is similar to the standard problem that if we deny regularity, will it be regular irregularity or irregular irregularity, that is, at what level are we denying the phenomeno? And only at one level?

[-]Richard_Ngo4y170

It seems to me that Eliezer overrates the concept of a simple core of general intelligence, whereas Paul underrates it. Or, alternatively: it feels like Eliezer is leaning too heavily on the example of humans, and Paul is leaning too heavily on evidence from existing ML systems which don't generalise very well.

I don't think this is a particularly insightful or novel view, but it seems worth explicitly highlighting that you don't have to side with one worldview or the other when evaluating the debates between them. (Although I'd caution not to just average their two views - instead, try to identify Eliezer's best arguments, and Paul's best arguments, and reconcile them.)

[-]Richard_Ngo1y160

(Vague, speculative thinking): Is the time element of UDT actually a distraction? Consider the following: agents A and B are in a situation where they'd benefit from cooperation. Unfortunately, the situation is complicated—it's not like a prisoner's dilemma, where there's a clear "cooperate" and a clear "defect" option. Instead they need to take long sequences of actions, and they each have many opportunities to subtly gain an advantage at the other's expense.

Therefore instead of agreements formulated as "if you do X I'll do Y", it'd be far more beneficial for them to make agreements of the form "if you follow the advice of person Z then I will too". Here person Z needs to be someone that both A and B trust to be highly moral, neutral, competent, etc. Even if there's some method of defecting that neither of them considered in advance, at the point in time when it arises Z will advise against doing it. (They don't need to actually have access to Z, they can just model what Z will say.)

If A and B don't have much communication bandwidth between them (e.g. they're trying to do acausal coordination) then they will need to choose a Z that's a clear Schelling point, even if that Z is subo... (read more)

2Vladimir_Nesov1y

Most discussion of updatelessness suggests that Z is an agent similar to A and B, and also that it's a policy whose implications are transparent to A and B. I think an importaint case has Z quite unlike A or B, possibly much smaller and more legible than them. And it can still be an agent in its own right, capable of eventually growing stronger than A or B were at the time Z was initially formulated. By a growing/developing Z I mean something that exists in coordination through both of these places, rather than splitting into a version of Z at A, and a version of Z at B, losing touch with each other. Such a Z might be thought of as an environmental agent that A and B create near them, equipped to keep in contact with its alternative instance (knowing enough about both its instance near A and its instance near B), rather than specifically a commitment of A or B, or a replacement of A or B, or a result of merging A and B. The commitment of A and B is then to future interactions with Z, which the updateless/coordinated core of Z should be sufficiently aware of to plan for.

1Martín Soto1y

I think Nesov had some similar idea about "agents deferring to a (logically) far-away algorithm-contract Z to avoid miscoordination", although I never understood it completely, nor think that idea can solve miscoordination in the abstract (only, possibly, be a nice pragmatic way to bootstrap coordination from agents who are already sufficiently nice). Hate to always be that guy, but if you are assuming all agents will only engage in symmetric commitments, then you are assuming commitment races away. In actuality, it is possible for a (meta-) commitment race to happen about "whether I only engage in symmetric commitments".

2Vladimir_Nesov1y

The central question to my mind is principles of establishing coordination between different situations/agents, and contracts is a framing for what coordination might look like once established. Agentic contracts have the additional benefit of maintaining coordination across their instances once it's established initially. Coordination theory should clarify how agents should think about establishing coordination with each other, how they should construct these contracts. This is not about niceness/cooperation. For example I think it should be possible to understand a transformer as being in coordination with the world through patterns in the world and circuits in the transformer, so that coordination gets established through learning. Beliefs are contracts between a mind and its object of study, essential tools the mind has for controlling it. Consequentialist control is a special case of coordination in this sense, and I think one problem with decision theories is that they are usually overly concerned with remaining close to consequentialist framing.

[-]Richard_Ngo2y162

[Epistemic status: rough speculation, feels novel to me, though Wei Dai probably already posted about it 15 years ago.]

UDT is (roughly) defined as "follow whatever commitments a past version of yourself would have made if they'd thought about your situation". But this means that any UDT agent is only as robust to adversarial attacks as their most vulnerable past self. Specifically, it creates an incentive for adversaries to show UDT agents situations that would trick their past selves into making unwise commitments. It also creates incentives for UDT agents themselves to hack their past selves, in order to artificially create commitments that "took effect" arbitrarily far back in their past.

In some sense, then, I think UDT might have a parallel structure to the overall alignment problem. You have dumber past agents who don't understand most of what's going on. You have smarter present agents who have trouble cooperating, because they know too much. The smarter agents may try to cooperate by punting to "Schelling point" dumb agents. (But this faces many of the standard problems of dumb agents making decisions—e.g. the commitments they make will probably be inconsistent or incoherent... (read more)

6Wei Dai1y

This seems substantially different from UDT, which does not really have or use a notion of "past version of yourself". For example imagine a variant of Counterfactual Mugging in which there is no preexisting agent, and instead Omega creates an agent from scratch after flipping the coin and gives it the decision problem. UDT is fine with this but "follow whatever commitments a past version of yourself would have made if they'd thought about your situation" wouldn't work. I recall that I described "exceptionless decision theory" or XDT as "do what my creator would want me to do", which seems closer to your idea. I don't think I followed up the idea beyond this, maybe because I realized that humans aren't running any formal decision theory, so "what my creator would want me to do" is ill defined. (Although one could say my interest in metaphilosophy is related to this, since what I would want an AI to do is to solve normative decision theory using correct philosophical reasoning, and then do what it recommends.) Anyway, the upshot is that I think you're exploring a decision theory approach that's pretty distinct from UDT so it's probably a good idea to call it something else. (However there may be something similar in the academic literature, or someone described something similar on LW that I'm not familiar with or forgot.)

2Richard_Ngo1y

My terminology here was sloppy, apologies. When I say "past versions of yourself" I am also including (as Nesov phrases it below) "the idealized past agent (which doesn't physically exist)". E.g. in the Counterfactual Mugging case you describe, I am thinking about precommitments that the hypothetical past version of yourself from before the coin was flipped would have committed to. I find it a more intuitive way to think about UDT, though I realize it's a somewhat different framing from yours. Do you still think this is substantially different?

4Vladimir_Nesov2y

UDT never got past the setting of unchanging preferences, so the present agent blindly defers to all decisions of the idealized past agent (which doesn't physically exist). And if the past agent doesn't try to wade in the murky waters of logical updatelessness, it's not really dumber or more fallible to trickery, it can see everything the way a universal Turing machine or Solomonoff induction can "see everything". Coordinating agents with different values was instead explored under the heading of Prisoner's Dilemma. Though a synthesis between coordination of agents with different values and UDT (recognizing Schelling point contracts as a central construction) is long overdue.

2Richard_Ngo2y

I actually think it might still be more fallible, for a couple of reasons. Firstly, consider an agent which, at time T, respects all commitments it would have made at times up to T. Now if you're trying to attack the agent at time T, you have T different versions of it that you can attack, and if any of them makes a dumb commitment then you win. I guess you could account for this by just gradually increasing the threshold for making commitments over time, though. Secondly: the further back you go, the more farsighted the past agent needs to be about the consequences of its commitments. If you have any compounding mistakes in the way it expects things to play out, then it'll just get worse and worse the further back you defer. Again, I guess you could account for this by having a higher threshold for making commitments which you expect to benefit you further down the line. Then, re logical updatelessness: it feels like in the long term we need to unify logical + empirical updates, because they're roughly the same type of thinking. Murky waters perhaps, but necessary ones. Yeah, so what could this look like? I think one important idea is that you don't have to be deferring to your past self, it's just that your past self is the clearest Schelling point. But it wouldn't be crazy for me to, say, use BDT: Buddha Decision Theory, in which I obey all commitments that the Buddha would have made for me if he'd been told about my situation. The differences between me using UDT and BDT (tentatively) seem only qualitative to me, not quantitative. BDT makes it harder for me to cooperate with hypothetical copies of myself who hadn't yet thought of BDT (because "Buddha" is less of a Schelling point amongst copies of myself than "past Richard"). It also makes me worse off than UDT in some cases, because sometimes the Buddha would make commitments in favor of his interests, not mine. But it also makes it a bit easier for me to cooperate with others, who might also converge to

4Vladimir_Nesov1y

UDT doesn't do multistage commitments, it has a single all-powerful "past" version that looks into all possible futures before pronouncing a global policy that all of them would then follow. This policy is not a collection of commitments in a reasonable informal sense, it's literally all details of behavior of future versions of the agent in response to all possible observations. In case of logical updatelessness, also in response to all possible observations of computational facts. (UDT for the idealized past version defines a single master model, future versions are just passively running inference from the contexts of their particular situations.) The convergent idea for acausal coordination between systems A and B seems to be constructing a shared subagent C whose instances exist as part of both A and B (after A and B successfully both construct the same C, not before), so that C can then act within them in the style of FDT, though really it's mostly about C thinking of the effects of its behavior in terms of "I am an algorithm" rather than "I am a physical object". (For UDT, the shared subagent C is the idealized common past version of its different possible future versions A and B. This assumes that A and B already have a lot in common, so maybe C is instead Buddha.) A bulk of the blind alleys seem to be about allowing subagents various superpowers, instead of focusing on managing the fallout of making them small and bounded (but possibly more plentiful). I think this is where investigations into logical updatelessness go wrong. It does need solving, but not by considering some fact unknown globally, or even at certain logical times. Instead a fact can remain unknown to some small subagent, and can be observed by it at some point, or computed by another subagent. Values are also knowledge, so sufficiently small subagents shouldn't even by default know full values of the larger system, and should be prepared to learn more about them. This is a consideration t

4Richard_Ngo1y

Okay, so trying to combine Prisoner's dilemma and UDT, we get: A and B are in a prisoner's dilemma. Suppose they have a list of N agents (which include, say, A's past self, B's past self, the Buddha, etc), and they each must commit to following one of those agent's instructions. Each of them estimates: "conditional on me committing to listen to agent K, here's a distribution over which agent they'd commit to listen to" And then you maximize expected value based on that. Okay, but why isn't this exactly the same as them just thinking to themselves "conditional on me taking action K, here's the distribution over their actions" for each of N actions they could take, and then maximizing expected value? It feels like the difference is that it's really hard to actually reason about the correlations between my low-level actions and your low-level actions, whereas it might be easier to reason about the correlations between my high-level commitments and your high-level commitments. I.e. the role of the Buddha in this situation is just to make the acausal coordination here much easier.

2Vladimir_Nesov1y

The main trick with PD is that instead of an agent only having two possible actions C and D, we consider many programs the agent might self-modify into (commit to becoming) that each might in the end compute C or D. This effectively changes the action space, there are now many more possible actions. And these programs/actions can be given access (like quines, by their own construction) to initial source code of all the agents, allowed to reason about them. But then programs have logical uncertainty about how they in the end behave, so the things you'd be enumerating don't immediately cash out in expected values. And these programs can decide to cause different expected values depending of what you'll do with their behavior, anticipate how you reason about them through reasoning about you in turn. It's hard to find clear arguments for why any particular desirable thing could happen as a result of this setup. UDT is notable for being one way of making this work. The "open source game theory" of PD (through Löb's theorem, modal fixpoints, Payor's lemma) pinpoints some cases where it's possible to say that we get cooperation in PD. But in general it's proven difficult to say anything both meaningful and flexible about this seemingly in-broad-strokes-inevitable setup, in particular for agents with different values that are doing more general things than playing PD. (The following relies a little bit on motivation given in the other comment.) When both A and B consider listening to a shared subagent C, subagent C is itself considering what it should be doing, depending on what A and B do with C's behavior. So for example with A there are two stages of computation to consider: first, it was A and didn't yet decide to sign the contract, then it became a composite system P(C), where P is A's policy for giving influence to C's behavior (possibly P and A include a larger part of the world where the first agent exists, not just the agent itself). The commitment of A is to th

[-]Richard_Ngo4y150

I've been reading Eliezer's recent stories with protagonists from dath ilan (his fictional utopia). Partly due to the style, I found myself bouncing off a lot of the interesting claims that he made (although it still helped give me a feel for his overall worldview). The part I found most useful was this page about the history of dath ilan, which can be read without much background context. I'm referring mostly to the exposition on the first 2/3 of the page, although the rest of the story from there is also interesting. One key quote from the remainder of the story:

"The next most critical fact about Earth is that from a dath ilani perspective their civilization is made entirely out of coordination failure. Coordination that fails on every scale recursively, where uncoordinated individuals assemble into groups that don't express their preferences, and then those groups also fail to coordinate with each other, forming governments that offend all of their component factions, which governments then close off their borders from other governments. The entirety of Earth is one gigantic failure fractal. It's so far below the multi-agent-optimal-boundary, only their profess

... (read more)

1AprilSR4y

I’d say lots of other things he’s said support that update. Stuff about how your model of the world will be accurate if and only if you somehow approximate Bayes’ law, for example. The dath ilan based fiction definitely helped me internalize the idea better though.

[-]Richard_Ngo2y122

A tension that keeps recurring when I think about philosophy is between the "view from nowhere" and the "view from somewhere", i.e. a third-person versus first-person perspective—especially when thinking about anthropics.

One version of the view from nowhere says that there's some "objective" way of assigning measure to universes (or people within those universes, or person-moments). You should expect to end up in different possible situations in proportion to how much measure your instances in those situations have. For example, UDASSA ascribes measure based on the simplicity of the computation that outputs your experience.

One version of the view from somewhere says that the way you assign measure across different instances should depend on your values. You should act as if you expect to end up in different possible future situations in proportion to how much power to implement your values the instances in each of those situations has. I'll call this the ADT approach, because that seems like the core insight of Anthropic Decision Theory. Wei Dai also discusses it here.

In some sense each of these views makes a prediction. UDASSA predicts that we live in a universe with laws of physi... (read more)

2Ben2y

Very interesting. It sounds like your "third person view from nowhere" vs the "first person view from somewhere" is very similar to something I was thinking about recently. I called them "objectively distinct situations" in contrast with "subjectively distinct situations". My view is that most of the anthropic arguments that "feel wrong" to me are built on trying to make me assign equal probability to all subjectively distinct scenarios, rather than objective ones. eg. A replication machine makes it so there are two of me, then "I" could be either of them, leaving two subjectively distinct cases, even if on the object level there is actual no distinction between "me" being clone A or clone B. [1] I am very sceptical of this ADT. If you think the time/place you have ended up is unusually important I think that is more likely explained by something like "people decide what is important based on what is going on around them". [1] My thoughts are here: https://www.lesswrong.com/posts/v9mdyNBfEE8tsTNLb/subjective-questions-require-subjective-information

2Wei Dai2y

I'm not sure this is a valid interpretation of ADT. Can you say more about why you interpret ADT this way, maybe with an example? My own interpretation of how UDT deals with anthropics (and I'm assuming ADT is similar) is "Don't think about indexical probabilities or subjective anticipation. Just think about measures of things you (considered as an algorithm with certain inputs) have influence over." This seems to "work" but anthropics still feels mysterious, i.e., we want an explanation of "why are we who we are / where we're at" and it's unsatisfying to "just don't think about it". UDASSA does give an explanation of that (but is also unsatisfying because it doesn't deal with anticipations, and also is disconnected from decision theory). I would say that under UDASSA, it's perhaps not super surprising to be when/where we are, because this seems likely to be a highly simulated time/scenario for a number of reasons (curiosity about ancestors, acausal games, getting philosophical ideas from other civilizations).

2Richard_Ngo2y

(Speculative paragraph, quite plausibly this is just nonsense.) Suppose you have copies A and B who are both offered the same bet on whether they're A. One way you could make this decision is to assign measure to A and B, then figure out what the marginal utility of money is for each of A and B, then maximize measure-weighted utility. Another way you could make this decision, though, is just to say "the indexical probability I assign to ending up as each of A and B is proportional to their marginal utility of money" and then maximize your expected money. Intuitively this feels super weird and unjustified, but it does make the "prediction" that we'd find ourselves in a place with high marginal utility of money, as we currently do. (Of course "money" is not crucial here, you could have the same bet with "time" or any other resource that can be compared across worlds.) Fair point. By "acausal games" do you mean a generalization of acausal trade? (Acausal trade is the main reason I'd expect us to be simulated a lot.)

2Wei Dai2y

This is particularly weird because your indexical probability then depends on what kind of bet you're offered. In other words, our marginal utility of money differs from our marginal utility of other things, and which one do you use to set your indexical probability? So this seems like a non-starter to me... (ETA: Maybe it changes moment by moment as we consider different decisions, or something like that? But what about when we're just contemplating a philosophical problem and not trying to make any specific decisions?) Yes, didn't want to just say "acausal trade" in case threats/war is also a big thing.

2Richard_Ngo2y

It seems pretty weird to me too, but to steelman: why shouldn't it depend on the type of bet you're offered? Your indexical probabilities can depend on any other type of observation you have when you open your eyes. E.g. maybe you see blue carpets, and you know that world A is 2x more likely to have blue carpets. And hearing someone say "and the bet is denominated in money not time" could maybe update you in an analogous way. I mostly offer this in the spirit of "here's the only way I can see to reconcile subjective anticipation with UDT at all", not "here's something which makes any sense mechanistically or which I can justify on intuitive grounds".

6Wei Dai2y

I added this to my comment just before I saw your reply: Maybe it changes moment by moment as we consider different decisions, or something like that? But what about when we're just contemplating a philosophical problem and not trying to make any specific decisions? Ah I see. I think this is incomplete even for that purpose, because "subjective anticipation" to me also includes "I currently see X, what should I expect to see in the future?" and not just "What should I expect to see, unconditionally?" (See the link earlier about UDASSA not dealing with subjective anticipation.) ETA: Currently I'm basically thinking: use UDT for making decisions, use UDASSA for unconditional subjective anticipation, am confused about conditional subjective anticipation as well as how UDT and UDASSA are disconnected from each other (i.e., the subjective anticipation from UDASSA not feeding into decision making). Would love to improve upon this, but your idea currently feels worse than this...

[-]Richard_Ngo5y120

In a bayesian rationalist view of the world, we assign probabilities to statements based on how likely we think they are to be true. But truth is a matter of degree, as Asimov points out. In other words, all models are wrong, but some are less wrong than others.

Consider, for example, the claim that evolution selects for reproductive fitness. Well, this is mostly true, but there's also sometimes group selection, and the claim doesn't distinguish between a gene-level view and an individual-level view, and so on...

So just assigning it a single probability seems inadequate. Instead, we could assign a probability distribution over its degree of correctness. But because degree of correctness is such a fuzzy concept, it'd be pretty hard to connect this distribution back to observations.

Or perhaps the distinction between truth and falsehood is sufficiently clear-cut in most everyday situations for this not to be a problem. But questions about complex systems (including, say, human thoughts and emotions) are messy enough that I expect the difference between "mostly true" and "entirely true" to often be significant.

Has this been discussed before? Given Less Wrong's name, I'd be surprised if not, but I don't think I've stumbled across it.

8habryka5y

This feels generally related to the problems covered in Scott and Abram's research over the past few years. One of the sentences that stuck out to me the most was (roughly paraphrased since I don't want to look it up): I.e. our current formulations of bayesianism like solomonoff induction only formulate the idea of a hypothesis at such a low level that even trying to think about a single hypothesis rigorously is basically impossible with bounded computational time. So in order to actually think about anything you have to somehow move beyond naive bayesianism.

2Richard_Ngo5y

This seems reasonable, thanks. But I note that "in order to actually think about anything you have to somehow move beyond naive bayesianism" is a very strong criticism. Does this invalidate everything that has been said about using naive bayesianism in the real world? E.g. every instance where Eliezer says "be bayesian". One possible answer is "no, because logical induction fixes the problem". My uninformed guess is that this doesn't work because there are comparable problems with applying to the real world. But if this is your answer, follow-up question: before we knew about logical induction, were the injunctions to "be bayesian" justified? (Also, for historical reasons, I'd be interested in knowing when you started believing this.)

2habryka5y

I think it definitely changed a bunch of stuff for me, and does at least a bit invalidate some of the things that Eliezer said, though not actually very much. In most of his writing Eliezer used bayesianism as an ideal that was obviously unachievable, but that still gives you a rough sense of what the actual limits of cognition are, and rules out a bunch of methods of cognition as being clearly in conflict with that theoretical ideal. I did definitely get confused for a while and tried to apply Bayes to everything directly, and then felt bad when I couldn't actually apply bayes theorem in some situations, which I now realize is because those tended to be problems where embededness or logical uncertainty mattered a lot. My shift on this happened over the last 2-3 years or so. I think starting with Embedded Agency, but maybe a bit before that.

2Richard_Ngo5y

Which ones? In Against Strong Bayesianism I give a long list of methods of cognition that are clearly in conflict with the theoretical ideal, but in practice are obviously fine. So I'm not sure how we distinguish what's ruled out from what isn't. Can you give an example of a real-world problem where logical uncertainty doesn't matter a lot, given that without logical uncertainty, we'd have solved all of mathematics and considered all the best possible theories in every other domain?

2habryka5y

I think in-practice there are lots of situations where you can confidently create a kind of pocket-universe where you can actually consider hypotheses in a bayesian way. Concrete example: Trying to figure out who voted a specific way on a LW post. You can condition pretty cleanly on vote-strength, and treat people's votes as roughly independent, so if you have guesses on how different people are likely to vote, it's pretty easy to create the odds ratios for basically all final karma + vote numbers and then make a final guess based on that. It's clear that there is some simplification going on here, by assigning static probabilities for people's vote behavior, treating them as independent (though modeling some subset of independence wouldn't be too hard), etc.. But overall I expect it to perform pretty well and to give you good answers. (Note, I haven't actually done this explicitly, but my guess is my brain is doing something pretty close to this when I do see vote numbers + karma numbers on a thread) Well, it's obvious that anything that claims to be better than the ideal bayesian update is clearly ruled out. I.e. arguments that by writing really good explanations of a phenomenon you can get to a perfect understanding. Or arguments that you can derive the rules of physics from first principles. There are also lots of hypotheticals where you do get to just use Bayes properly and then it provides very strong bounds on the ideal approach. There are a good number of implicit models behind lots of standard statistics models that when put into a bayesian framework give rise to a more general formulation. See the Wikipedia article for "Bayesian interpretations of regression" for a number of examples. Of course, in reality it is always unclear whether the assumptions that give rise to various regression methods actually hold, but I think you can totally say things like "given these assumption, the bayesian solution is the ideal one, and you can't perform better th

2Raemon5y

Are you able to give examples of the times you tried to be Bayesian and it failed because embedded was?

1EpicNamer270985y

Scott and Abram? Who? Do they have any books I can read to familiarize myself with this discourse?

2habryka5y

Scott: https://lesswrong.com/users/scott-garrabrant Abram: https://lesswrong.com/users/abramdemski

2Richard_Ngo5y

Scott Garrabrant and Abram Demski, two MIRI researchers. For introductions to their work, see the Embedded Agency sequence, the Consequences of Logical Induction sequence, and the Cartesian Frames sequence.

5DanielFilan5y

Related but not identical: this shortform post.

2Zack_M_Davis5y

See the section about scoring rules in the Technical Explanation.

2Richard_Ngo5y

Hmmm, but what does this give us? He talks about the difference between vague theories and technical theories, but then says that we can use a scoring rule to change the probabilities we assign to each type of theory. But my question is still: when you increase your credence in a vague theory, what are you increasing your credence about? That the theory is true? Nor can we say that it's about picking the "best theory" out of the ones we have, since different theories may overlap partially.

7Zack_M_Davis5y

If we can quantify how good a theory is at making accurate predictions (or rather, quantify a combination of accuracy and simplicity), that gives us a sense in which some theories are "better" (less wrong) than others, without needing theories to be "true".

[-]Richard_Ngo5yΩ6120

Oracle-genie-sovereign is a really useful distinction that I think I (and probably many others) have avoided using mainly because "genie" sounds unprofessional/unacademic. This is a real shame, and a good lesson for future terminology.

4adamShimi5y

After rereading the chapter in Superintelligence, it seems to me that "genie" captures something akin to act-based agents. Do you think that's the main way to use this concept in the current state of the field, or do you have other applications in mind?

2Richard_Ngo5y

Ah, yeah, that's a great point. Although I think act-based agents is a pretty bad name, since those agents may often carry out a whole bunch of acts in a row - in fact, I think that's what made me overlook the fact that it's pointing at the right concept. So not sure if I'm comfortable using it going forward, but thanks for pointing that out.

4DanielFilan5y

Perhaps the lesson is that terminology that is acceptable in one field (in this case philosophy) might not be suitable in another (in this case machine learning).

4Richard_Ngo5y

I don't think that even philosophers take the "genie" terminology very seriously. I think the more general lesson is something like: it's particularly important to spend your weirdness points wisely when you want others to copy you, because they may be less willing to spend weirdness points.

1adamShimi5y

Is that from Superintelligence? I googled it, and that was the most convincing result.

2Richard_Ngo5y

Yepp.

[-]Richard_Ngo9mo114

In my post on value systematization I used utilitarianism as a central example of value systematization.

Value systematization is important because it's a process by which a small number of goals end up shaping a huge amount of behavior. But there's another different way in which this happens: core emotional motivations formed during childhood (e.g. fear of death) often drive a huge amount of our behavior, in ways that are hard for us to notice.

Fear of death and utilitarianism are very different. The former is very visceral and deep-rooted; it typically influences our behavior via subtle channels that we don't even consciously notice (because we suppress a lot of our fears). The latter is very abstract and cerebral, and it typically influences our behavior via allowing us to explicitly reason about which strategies to adopt.

But fear of death does seem like a kind of value systematization. Before we have a concept of death we experience a bunch of stuff which is scary for reasons we don't understand. Then we learn about death, and then it seems like we systematize a lot of that scariness into "it's bad because you might die".

But it seems like this is happening way less consciously th... (read more)

2Thane Ruthenis9mo

I think that's right. Taking on the natural-abstraction lens, there is a "ground truth" to the "hierarchy of values". That ground truth can be uncovered either by "manual"/symbolic/System-2 reasoning, or by "automatic"/gradient-descent-like/System-1 updates, and both processes would converge to the same hierarchy. But in the System-2 case, the hierarchy would be clearly visible to the conscious mind, whereas the System-1 route would make it visible only indirectly, by the impulses you feel. I don't know about the conflict thing, though. Why do you think System 2 would necessarily oppose System 1's deepest motivations?

1Lucien9mo

Reminds me of Maslow's pyramid. I made an article about values, saying the supreme value is life and every other value derives from it. Watch out, this most probably does not align with your view at first glance: https://www.lesswrong.com/posts/xx3St4KC3KHHPGfL9/human-alignment

1robo9mo

I don't think it's system 1 doing the systemization. Evolution beat fear of death into us in lots of independent forms (fear of heights, snakes, thirst, suffocation, etc.), but for the same underlying reason. Fear of death is not just an abstraction humans invented or acquired in childhood; is a "natural idea" pointed at by our brain's innate circuitry from many directions. Utilitarianism doesn't come with that scaffolding. We don't learn to systematize Euclidian and Minkowskian spaces the same way either.

[-]Richard_Ngo2yΩ8115

People sometimes try to reason about the likelihood of deceptive alignment by appealing to speed priors and simplicity priors. I don't like such appeals, because I think that the differences between aligned and deceptive AGIs will likely be a very small proportion of the total space/time complexity of an AGI. More specifically:

1. If AGIs had to rederive deceptive alignment in every episode, that would make a big speed difference. But presumably, after thinking about it a few times during training, they will remember their conclusions for a while, and bring them to mind in whichever episodes they're relevant. So the speed cost of deception will be amortized across the (likely very long) training period.

2. AGIs will represent a huge number of beliefs and heuristics which inform their actions (e.g. every single fact they know). A heuristic like "when you see X, initiate the world takeover plan" would therefore constitute a very small proportion of the total information represented in the network; it'd be hard to regularize it away without regularizing away most of the AGI's knowledge.

I think that something like the speed vs simplicity tradeoff is relevant to the likelihood of deceptiv... (read more)

8ryan_greenblatt2y

Why do you think SGD will do this? Or are you imagining non-SGD mechanisms? It seems non-obvious to me that this will occur with SGD, though possible.

21a3orn2y

You mean this about something trained totally differently than a LLM, no? Because this mechanism seems totally implausible to me otherwise.

4the gears to ascension2y

Think during the forward pass, learn during the backward pass; if the model uses deceptive reasoning in the forward pass and the gradient says it's useful for prediction, that seems like the mechanism as described. Thoughts?

41a3orn2y

So, there are a few different reasons, none of which I've formalized to my satisfaction. I'm curious if these make sense to you. (1) One is that the actual kinds of reasoning that an LLM can learn in its forward pass are quite limited. As is well established, for instance, Transformers cannot multiply arbitrarily-long integers in a single forward pass. The number of additions involved in multiplying an N-digit integer increases in an unbounded way with N; thus, a Transformer with with a finite number of layers cannot do it. (Example: Prompt GPT-4 for the results of multiplying two 5-digit numbers, specifying not to use a calculator, see how it does.) Of course in use you can teach a GPT to use a calculator -- but we're talking about operations that occur in single forward pass, which rules out using tools. Because of this shallow serial depth, a Transformer also cannot (1) divide arbitrary integers, (2) figure out the results of physical phenomena that have multiplication / division problems embedded in them, (3) figure out the results of arbitrary programs with loops, and so on. (Note -- to be very clear NONE of this is a limitation on what kind of operations we can get a transformer to do over multiple unrollings of the forward pass. You can teach a transformer to use a calculator; or to ask a friend for help; or to use a scratchpad, or whatever. But we need to hide deception in a single forward pass, which is why I'm harping on this.) So to think that you learn deception in forward pass, you have to think that the transformer thinks something like "Hey, if I deceive the user into thinking that I'm a good entity, I'll be able to later seize power, and if I seize power, then I'll be able to (do whatever), so -- considering all this, I should... predict the next token will be "purple"" -- and that it thinks this in a context that could NOT come up with the algorithm for multiplication, or for addition, or for any number of other things, even though an algorith

4Daniel Kokotajlo2y

You've given me a lot to think about, thanks! Here are my thoughts as I read: Yes, and the sort of deceptive reasoning I'm worried about sure seems pretty simple, very little serial depth to it. Unlike multiplying two 5-digit integers. For example the example you give involves like 6 steps. I'm pretty sure GPT4 already does 'reasoning' of about that level of sophistication in a single forward pass, e.g. when predicting the next token of the transcript of a conversation in which one human is deceiving another about something. (In fact, in general, how do you explain how LLMs can predict deceptive text, if they simply don't have enough layers to do all the deceptive reasoning without 'spilling the beans' into the token stream?) The reason I'm not worried about image segmentation models is that it doesn't seem like they'd have the relevant capabilities or goals. Maybe in the limit they would -- if we somehow banned all other kinds of AI, but let image segmentation models scale to arbitrary size and training data amounts -- then eventually after decades of scaling and adding 9's of reliability to their image predictions, they'd end up with scary agents inside because that would be useful for getting one of those 9's. But yeah, it's a pretty good bet that the relevant kinds of capabilities (e.g. ability to coherently pursue goals, ability to write code, ability to persuade humans of stuff) are most likely to appear earliest in systems that are trained in environments more tailored to developing those capabilities. tl;dr my answer is 'wake me when an image segmentation model starts performing well in dangerous capabilities evals like METR's and OpenAI's. Which won't happen for a long time because image segmentation models are going to be worse at agency than models explicitly trained to be agents." I think this is a misunderstanding of the LLM will learn deception hypothesis. First of all, the conditions of the hypothesis are not just "so long as it's big enough and

2the gears to ascension2y

These all seem like reasonable reasons to doubt the hypothesized mechanism, yup. I think you're underestimating how much can happen in a single forward pass, though - it has to be somewhat shallow, so it can't involve too many variables, but the whole point of making the networks as large as we do these days is that it turns out an awful lot can happen in parallel. I also think there would be no reason for deception to occur if it's never a good weight pattern to use to predict the data, it's only if the data contains a pattern that the gradient will put into a deceptive forward mechanism that this could possibly occur. For example, if the model is trained on a bunch of humans being deceptive about their political intentions, and then RLHF is attempted. In any case, I don't think the old yudkowsky model of deceptive alignment is relevant, in that I think the level of deception to expect from ai should be calibrated to be around the amount you'd expect from a young human, not some super schemer god. The concern arises only when the data actually contains patterns well modeled by deception, and this would be expected to be more present in the case of something like an engagement maximizer online learning RL system. And to be clear I don't expect the things that can destroy humanity to arise because of deception directly. It seems much more likely to me that they'll arise because competing people ask their model to do something that puts those models in competition in a way that puts humanity at risk, eg several different powerful competing model based engagement/sales optimizing reinforcement learners, or more speculatively something military. Something where the core problem is that alignment tech is effectively not used, and where solving this deception problem wouldn't have saved us anyway. Regarding the details of your descriptions: I really mainly think this sort of deception would arise in the wild when there's a reward model passing gradients to multiple ste

[-]Richard_Ngo2y90

It's really weird that we find ourselves at the hinge of history. One proposed explanation is that we're part of an ancestor simulation. It makes sense that ancestor simulations would be focused on the hinge of history. But unless ancestor simulations make up a significant proportion of future minds, it's still weird that we find ourselves in a simulation rather than actually experiencing the future.

Why might ancestor simulations make up a significant proportion of future minds? One possible answer is that ancestor simulations provide the information requi... (read more)

5mako yass2y

I've done work in this area, but never been particularly enthusiastic about promoting it. It usually turns out to be inactionable/grim/likely to rouse a panic. This is a familiar thought, to me. A counterargument occurs to me: Isn't it arguable that most of what we need to know about a species, to trade with it, is just downstream of its biology? Of course we talk a lot about our contingent factors, our culture, our history, but I think we're pretty much just the same animals we've always been, extrapolated. If that's the case, wouldn't far more simulation time be given to evolutionary histories, rather than than simulating variations of hinges? Anthropic measure wouldn't be especially concentrated on the hinge, it might even skip it. Countercounterargument: it also seems like there are a lot of anti-inductive effects in the histories of technological societies that might mean you really do have to simulate it all to find out how values settle or just to figure out the species' rate of success. Evolutionary histories might also have a lot more computationally compressible shared structure. I'd be surprised if this, the world in front of us, were a pareto-efficient bargaining outcome. Hinge histories fucking suck to live in and I would strongly prefer a trade protocol that instantiated as few of them as possible. I wouldn't expect many to be necessary, certainly not enough to significantly outweigh the... thing that is supposed to come after. (at this point, I'd prefer to take it into DMs/call) Thinking about this stuff again, something occurred to me. Please make sure to keep, in cold storage, copies of misaligned AGIs that you may produce, when you catch them. It's important. This policy could save us.

3Carl Feynman2y

Would you care to expand on your remark? I don’t see how it follows from what you said above it.

2mako yass2y

Yeah, it wasn't argued. I wasn't sure whether it needed to be explained, for Richard. I don't remember how I wound up getting there from the rest of the comment, I think it was just in the same broad neighborhood. Regardless, yes, I totally can expand on that. Here, I wrote it up: Do Not Delete your Misaligned AGI.

4FlorianH2y

World champion in Chess: "It's really weird that I'm world champion. It must be a simulation or I must dream or.." Joe Biden: "It's really weird I'm president, it must be a simul..." (Donald Trump: "It really really makes no sense I'm president, it MUST be a s..") David Chalmers: "It's really weird I'm providing the seminal hard problem formulation. It must be a sim.." ... Rationalist (before finding lesswrong): "Gosh, all these people around me, really wired differently than I am. I must be in a simulation." Something seems funny to me in the anthropic reasoning in these examples, and in yours too. Of course we have one world champion in chess or anything, so a reasoning that means that world champion quasi by definition question's his champion-ness, seems odd. Then, I'd be lying if I claimed I could not intuitively empathize with his wondering about the odds of exactly him being the world champion among 9 billions. This leads me to the following, that eventually +- satisfies me: Hypothetically, imagine each generation has only 1 person, and there's rebirth: it's just a rebirth of the same person, in a different generation. With some simplification: 1. For 10 000 generations you lived in stone-age conditions 2. For 1 generation - today - you're the hinge-of-history generation 3. X (X being: you won't live anymore at all as AI killed everything; or you live 1 mio generations happily, served by AI, or what have you). The 10 000 you's didn't have much reason to wonder about hinge of history, and so doesn't happen to think about it. The one you, in the hinge-of-history generation, by definition, has much reasons to think about the hinge-of-history, and does think about it. So, it has becomes a bit like a lottery game, which you repeat so many times until you naturally once draw the winning number. At that lucky punch, there's no reason to think "Unlikely, it's probably a simulation", or anything. I have the impression in the similar way, the reinc

2Daniel Kokotajlo2y

Nit: ECL is just one of several kinds of acausal cooperation across large worlds.

2Richard_Ngo2y

What are the others?

0JBlack2y

In general I don't think anthropic reasoning like this holds any substance. We experience what we experience, and condition on that in forming models about what it is and where we are in it. We don't get to make millions of bits of observations about being a human in a technological society, use those observations to extrapolate the possibility of supergalactic multitudes of consciousness, and then express surprise at a pathetic few dozen bits of improbability of not being one of those multitudes. We already used those bits (and a great many more!) in forming our model in the first place.

[-]Richard_Ngo2y80

Since there's been some recent discussion of the SSC/NYT incident (in particular via Zack's post), it seems worth copying over my twitter threads from that time about why I was disappointed by the rationalist community's response to the situation.

I continue to stand by everything I said below.

Thread 1 (6/23/20):

Scott Alexander is the most politically charitable person I know. Him being driven off the internet is terrible. Separately, it is also terrible if we have totally failed to internalize his lessons, and immediately leap to the conclusion that the NY... (read more)

9Jiro2y

Scott is already too charitable. I'd even say that Scott being too charitable made this specific situation worse. I don't find this to be a worthwhile thing about Scott either for us to emulate, or for Scott to take further. "Quokka" is a meme about rationalists for a reason. You are not going to have unerring logical evidence that someone wants to harm you if they are trying to be at all subtle. You have to figure it out from their behavior. Sometimes it just isn't true that both sides are reasonable and have useful perspectives.

6Elizabeth2y

I think this only holds if NYT has a consistent policy of using real names. My understanding is they have repeatedly written about other people using pseudonyms only, and have not articulated a principled reason to treat Scott differently.

2Vladimir_Nesov2y

Scott's flavor of charity is not quite this. It wouldn't be useful for understanding sides that are not reasonable or have useless perspectives otherwise, or else you'd need to routinely "assume" false things to carry out the exercise. The point is to meaningfully engage with other perspectives, without the usual prerequisite of having positive beliefs about them. Treating them in a similar way as if they were reasonable or useful, even when they clearly aren't. Sometimes the resulting investigation changes one's mind on this point. But often it doesn't, while still revealing many details that wouldn't otherwise be noticed. Actually intervening on your own beliefs would be self-deception, while treating useless and unreasonable views as they are usually treated wouldn't be charity. This is related to tolerance, where the point isn't to start liking people you don't like, or to start considering them part of your own ingroup. It's instead an intervention/norm that goes around the dislike to remove some of its downsides without directly removing the dislike itself.

[-]Richard_Ngo3y80

My mental one-sentence summary of how to think about ELK is "making debate work well in a setting where debaters are able to cite evidence gained by using interpretability tools on each other".

I'm not claiming that this is how anyone else thinks about ELK (although I got the core idea from talking to Paul) but since I haven't seen it posted online yet, and since ELK is pretty confusing, I thought it'd be useful to put out there. In particular, this framing motivates us generating interpretability tools which scale in the sense of being robust when used as ... (read more)

[-]Richard_Ngo4y80

Being nice because you're altruistic, and being even nicer for decision-theoretic reasons on top of that, seems like it involves some kind of double-counting: the reason you're altruistic in the first place is because evolution ingrained the decision theory into your values.

But it's not fully double-counting: many humans generalise altruism in a way which leads them to "cooperate" far more than is decision-theoretically rational for the selfish parts of them - e.g. by making big sacrifices for animals, future people, etc. I guess this could be selfishly ra... (read more)

4Dagon4y

Your actions and decisions are not doubled. If you have multiple paths to arrive at the same behaviors, that doesn't make them wrong or double-counted, it just makes it hard to tell which of them is causal (aka: your behavior is overdetermined). Are you using "updatelessness" to refer to not having self in your utility function? If so, that's a new one one me, and I'd prefer "altruism" as the term. I'm not sure that the decision-theory use of "updateless" (to avoid incorrect predictions where experience is correlated with the question at hand) makes sense here.

2Richard_Ngo4y

Oh, this also suggests a way in which the utility function abstraction is leaky, because the reasons for the payoffs in a game may matter. E.g. if one payoff is high because the corresponding agent is altruistic, then in some sense that agent is "already cooperating" in a way which is baked into the game, and so the rational thing for them to do might be different from the rational thing for another agent who gets the same payoffs, but for "selfish" reasons. Maybe FDT already lumps this effect into the "how correlated are decisions" bucket? Idk.

[-]Richard_Ngo2y*70

In UDT2, when you're in epistemic state Y and you need to make a decision based on some utility function U, you do the following:
1. Go back to some previous epistemic state X and an EDT policy (the combination of which I'll call the non-updated agent).
2. Spend a small amount of time trying to find the policy P which maximizes U based on your current expectations X.
3. Run P(Y) to make the choice which maximizes U.

The non-updated agent gets much less information than you currently have, and also gets much less time to think. But it does use the same utility ... (read more)

5Martín Soto2y

People back then certainly didn't think of changing preferences. Also, you can get rid of this problem by saying "you just want to maximize the variable U". And the things you actually care about (dogs, apples) are just "instrumentally" useful in giving you U. So for example, it is possible in the future you will learn dogs give you a lot of U, or alternatively that apples give you a lot of U. Needless to say, this "instrumentalization" of moral deliberation is not how real agents work. And leads to getting Pascal's mugged by the world in which you care a lot about easy things. It's more natural to model U as a logically uncertain variable, freely floating inside your logical inductor, shaped by its arbitrary aesthetic preferences. This doesn't completely miss the importance of reward in shaping your values, but it's certainly very different to how frugally computable agents do it. I simply think the EV maximization framework breaks here. It is a useful abstraction when you already have a rigid enough notion of value, and are applying these EV calculations to a very concrete magisterium about which you can have well-defined estimates. Otherwise you get mugged everywhere. And that's not how real agents behave.

2Richard_Ngo2y

But you need some mechanism for actually updating your beliefs about U, because you can't empirically observe U. That's the role of V. I think this is fine. Consider two worlds: In world L, lollipops are easy to make, and paperclips are hard to make. In world P, it's the reverse. Suppose you're a paperclip-maximizer in world L. And a lollipop-maximizer comes up to you and says "hey, before I found out whether we were in L or P, I committed to giving all my resources to paperclip-maximizers if we were in P, as long as they gave me all their resources if we were in L. Pay up." UDT says to pay here—but that seems basically equivalent to getting "mugged" by worlds where you care about easy things.

1Martín Soto2y

Yep, but you can just treat it as another observation channel into UDT. You could, if you want, treat it as a computed number you observe in the corner of your eye, and then just apply UDT maximizing U, and you don't need to change UDT in any way. (Let's not forget this depends on your prior, and we don't have any privileged way to assign priors to these things. But that's a tangential point.) I do agree that there's not any sharp distinction between situations where it "seems good" and situations where it "seems bad" to get mugged. After all, if all you care about is maximizing EV, then you should take all muggings. It's just that, when we do that, something feels off (to us humans, maybe due to risk-aversion), and we go "hmm, probably this framework is not modelling everything we want, or missing some important robustness considerations, or whatever, because I don't really feel like spending all my resources and creating a lot of disvalue just because in the world where 1 + 1 = 3 someone is offering me a good deal". You start to see how your abstractions might break, and how you can't get any satisfying notion of "complete updatelessness" (that doesn't go against important intuitions). And you start to rethink whether this is what we normatively want, nor what we realistically see in agents.

2Richard_Ngo2y

Hmm, I'm confused by this. Why should we treat it this way? There's no actual observation channel, and in order to derive information about utilities from our experiences, we need to specify some value learning algorithm. That's the role V is playing. Obviously I am not arguing that you should agree to all moral muggings. If a pain-maximizer came up to you and said "hey, looks like we're in a world where pain is way easier to create than pleasure, give me all your resources", it would be nuts to agree, just like it would be nuts to get mugged by "1+1=3". I'm just saying that "sometimes you get mugged" is not a good argument against my position, and definitely doesn't imply "you get mugged everywhere".

1Martín Soto2y

Yes, absolutely! I just meant that, once you give me whatever V you choose to derive U from observations, I will just be able to apply UDT on top of that. So under this framework there doesn't seem to be anything new going on, because you are just choosing an algorithm V at the start of time, and then treating its outputs as observations. That's, again, why this only feels like a good model of "completely crystallized rigid values", and not of "organically building them up slowly, while my concepts and planner module also evolve, etc.".[1] Wait, but how does your proposal differ from EV maximization (with moral uncertainty as part of the EV maximization itself, as I explain above)? Because anything that is doing pure EV maximization "gets mugged everywhere". Meaning if you actually have the beliefs (for example, that the world where suffering is hard to produce could exist), you just take those bets. Of course if you don't have such "extreme" beliefs it doesn't, but then we're not talking about decision-making, and instead belief-formation. You could say "I will just do EV maximization, but never have extreme beliefs that lead to suspiciously-looking behavior", but that'd be hiding the problem under belief-formation, and doesn't seem to be the kind of efficient mechanism that agents really implement to avoid these failure modes. 1. ^ To be clear, V can be a very general algorithm (like "run a copy of me thinking about ethics"), so that this doesn't "feel like" having rigid values. Then I just think you're carving reality at the wrong spot. You're ignoring the actual dynamics of messy value formation, hiding them under V.

4Vladimir_Nesov2y

In times of UDT2, the background assumption was that agents should maintain an unchanging preference, which is separate from knowledge. One motivation for UDT is that updating makes an agent stop caring about updated-away possibilities, while UDT is not doing that. Going back to a previous epistemic state is a way of preserving preference from that epistemic state, the "current" utility function is considered a bug and doesn't do anything if UDT is adopted. The non-updated agent can in principle consider the information you currently have as one of the possibilities when formulating the general policy for all possibilities, though being bounded it won't do a very good job. Traditionally UDT1.1 wants to make its decisions from very little knowledge and to apply the policy to all always. A more pragmatic thing is to make decisions from modestly less knowledge and to scope the policy for middle-term future. Some form of this is useful for many thought experiments where the environment or other players also have the little knowledge our agent uses to make its decisions from the past, and so could know the policy the agent decides on before they need to prepare for it or make predictions about it. The problem is commitment races (as in the game of chicken), where everyone wants to decide earlier and force the others to respond. But there is a need to remain bounded in making decisions, both to personally compute them and to make it possible for others to anticipate them and to coordinate. This creates a more reasonable equilibrium, motivating decisions from a less ignorant epistemic state that have a better chance of being relevant to the current situation, in balance with trying to decide from a more ignorant epistemic state where a general policy would enable more strategicness across possibilities. UDT1.1 can't find such balance, but it's possible that something UDT2-shaped might.

2Richard_Ngo2y

I think there's an ambiguity here. UDT makes the agent stop considering updated-away possibilities, but I haven't seen any discussion of UDT which suggests that it stops caring about them in principle (except for a brief suggestion from Paul that one option for UDT is to "go back to a position where I’m mostly ignorant about the content of my values"). Rather, when I've seen UDT discussed, it focuses on updating or un-updating your epistemic state. I don't think the shift I'm proposing is particularly important, but I do think the idea that "you have your prior and your utility function from the very beginning" is a kinda misleading frame to be in, so I'm trying to nudge a little away from that.

2Vladimir_Nesov2y

UDT specifically enables agents to consider the updated-away possibilities in a way relevant to decision making, while an updated agent (that's not using something UDT-like) wouldn't be able to do that in any circumstance, and so would be functionally indistinguishable from an agent that has different preferences or undefined preferences for those possibilities. Not caring about them seems like an apt informal description (even as this is compatible with keeping the same utility function outside the event of current knowledge). In a similar way, we could say that after updating, an agent either changes their probability distribution or keeps the original prior. Historically it was overwhelmingly the frame until recently, so it's the correct frame for interpreting the intended meaning of texts from that time. This is a simplifying assumption that still leaves many open questions about how to make decisions in sufficiently strange situations (where merely models of behavior make these strange situations ubiquitous in practice). When an agent doesn't know its own preference and needs to do something about that, it's an additional complication that usually wasn't introduced.

2Richard_Ngo2y

Agreed; apologies for the sloppy phrasing. I agree, that's why I'm trying to outline an alternative frame for thinking about it.

2Richard_Ngo2y

Some more thoughts: we can portray the process of choosing a successor policy as the iterative process of making more and more commitments over time. But what does it actually look like to make a commitment? Well, consider an agent that is made of multiple subagents, that each get to vote on its decisions. You can think of a commitment as basically saying "this subagent still gets to vote, but no longer gets updated"—i.e. it's a kind of stop-gradient. Two interesting implications of this perspective: 1. The "cost" of a commitment can be measured both in terms of "how often does the subagent vote in stupid ways?", and also "how much space does it require to continue storing this subagent?" But since we're assuming that agents get much smarter over time, probably the latter is pretty small. 2. There's a striking similarity to the problem of trapped priors in human psychology. Parts of our brains basically are subagents that still get to vote but no longer get updated. And I don't think this is just a bug—it's also a feature. This is true on the level of biological evolution (you need to have a strong fear of death in order to actually survive) and also on the level of cultural evolution (if you can indoctrinate kids in a way that sticks, then your culture is much more likely to persist). The (somewhat provocative) way of phrasing this is that trauma is evolution's approach to implementing UDT. Someone who's been traumatized into conformity by society when they were young will then (in theory) continue obeying society's dictates even when they later have more options. Someone who gets very angry if mistreated in a certain way is much harder to mistreat in that way. And of course trauma is deeply suboptimal in a bunch of ways, but so too are UDT commitments, because they were made too early to figure out better alternatives. This is clearly only a small component of the story but the analogy is definitely a very interesting one.

5Richard_Ngo2y

More thoughts: what's the difference between paying in a counterfactual mugging based on: 1. Whether the millionth digit of pi (5) is odd or even 2. Whether or not there are an infinite number of primes? In the latter case knowing the truth is (near-)inextrictably entangled with a bunch of other capabilities, like the ability to do advanced mathematics. Whereas in the former it isn't. Suppose that before you knew either fact you were told that one of them was entangled in this way—would you still want to commit to paying out in a mugging based on it? Well... maybe? But it means that the counterlogical of "if there hadn't been an infinite number of primes" is not very well-defined—it's hard to modify your brain to add that belief without making a bunch of other modifications. So now Omega doesn't just have to be (near-)omniscient, it also needs to have a clear definition of the counterlogical that's "fair" according to your standards; without knowing that it has that, paying up becomes less tempting.

3Vladimir_Nesov2y

Individually logical counterfactuals don't seem very coherent. This is related to the "I'm an algorithm" vs. "I'm a physical object" distinction of FDT. When you are an algorithm considering a decision, you want to mark all sites of intervention/influence in the world where the world depends on your behavior. If you only mark some of them, then you later fail at the step where you ask what happens if you act differently, you obtain a broken counterfactual world where only some instances of the fact of your behavior have been replaced and not others. So I think it makes a bit more sense to ask where specifically your brain depends on a fact, to construct an exhausive dependence of your brain on the fact, before turning to particular counterfactual content for that fact to be replaced with. That is, dependence of a system on a fact, the way it varies with the fact, seems potentially clearer than individual counterfactuals of how that system works if the fact is set to be a certain way. (To make a somewhat hopeless analogy, fibration instead of individual fibers, and it shouldn't be a problem that all fibers are different from each other. Any question about a counterfactual should be reformulated into a question about a dependence.)

[-]Richard_Ngo3y*70

Random question I’ve been thinking about: how would you set up a market for votes? Suppose specifically that you have a proportional chances election (i.e. the outcome gets chosen with probability proportional to the number of votes cast for it—assume each vote is a distribution over candidates). So everyone has an incentive to get everyone who’s not already voting for their favorite option to change their vote; and you can have positive-sum trades where I sell you a promise to switch X% of my votes to a compromise candidate in exchange for you switching Y... (read more)

4Measure3y

Just spitballing here: Assign each voter 100 shares for each candidate. To vote, each voter selects a subset of their shares to constitute their vote. Voters can freely trade shares. Under this system, a voter would more highly value shares for candidates that are either very high or very low in their preference order (the later so as to exclude them from the vote). Thus, trades would look like each party exchanging shares about which they are themselves ambivalent to gain shares that are more valuable to them. If you remove the proportional chances part, then it becomes a guessing game of which marginal votes actually matter.

2Richard_Ngo3y

Interesting! Hadn't thought of this approach. Let's see... Intuitively I think it gets pretty strategically weird because a) who you vote for depends pretty sensitively on other peoples' votes (e.g. in proportional chances voting you want to vote for everyone who's above the expected value of everyone else's votes; in approval voting you want to vote for everyone you approve of unless it bumps them above someone you like more), and b) you want to buy from your enemies much more than from your friends, because your friends will already not be voting for bad candidates. But maybe the latter is fine because if you buy from your friends they'll end up with more money which they can then spend on other things? I'll keep thinking.

[-]Richard_Ngo4y*Ω570

I expect it to be difficult to generate adversarial inputs which will fool a deceptively aligned AI. One proposed strategy for doing so is relaxed adversarial training, where the adversary can modify internal weights. But this seems like it will require a lot of progress on interpretability. An alternative strategy, which I haven't yet seen any discussion of, is to allow the adversary to do a data poisoning attack before generating adversarial inputs - i.e. the adversary gets to specify inputs and losses for a given number of SGD steps, and then the adversarial input which the base model will be evaluated on afterwards. (Edit: probably a better name for this is adversarial meta-learning.)

[-]Richard_Ngo4y60

Another thought on dath ilan: notice how much of the work of Keltham's reasoning is based on him pattern-matching to tropes from dath ilani literature, and then trying to evaluate their respective probabilities. In other words: like bayesianism, he's mostly glossing over the "hypothesis generation" step of reasoning.

I wonder if dath ilan puts a lot of effort into spreading a wide range of tropes because they don't know how to teach systematically good hypothesis generation.

4Gunnar_Zarncke4y

I think you are overgeneralizing. We also see some mix of Dath Ilan, stories about Dath Ilan, stories about stories about Dath Ilan, and interactions between these, so all bets are off really.

[-]Richard_Ngo5yΩ460

I suspect that AIXI is misleading to think about in large part because it lacks reusable parameters - instead it just memorises all inputs it's seen so far. Which means the setup doesn't have episodes, or a training/deployment distinction; nor is any behaviour actually "reinforced".

4DanielFilan5y

I kind of think the lack of episodes makes it more realistic for many problems, but admittedly not for simulated games. Also, presumably many of the component Turing machines have reusable parameters and reinforce behaviour, altho this is hidden by the formalism. [EDIT: I retract the second sentence]

2DanielFilan5y

Actually I think this is total nonsense produced by me forgetting the difference between AIXI and Solomonoff induction.

2Richard_Ngo5y

Wait, really? I thought it made sense (although I'd contend that most people don't think about AIXI in terms of those TMs reinforcing hypotheses, which is the point I'm making). What's incorrect about it?

2DanielFilan5y

Well now I'm less sure that it's incorrect. I was originally imagining that like in Solomonoff induction, the TMs basically directly controlled AIXI's actions, but that's not right: there's an expectimax. And if the TMs reinforce actions by shaping the rewards, in the AIXI formalism you learn that immediately and throw out those TMs.

2Richard_Ngo5y

Oh, actually, you're right (that you were wrong). I think I made the same mistake in my previous comment. Good catch.

2[comment deleted]5y

4Steven Byrnes5y

Humans don't have a training / deployment distinction either... Do humans have "reusable parameters"? Not quite sure what you mean by that.

6Richard_Ngo5y

Yes we do: training is our evolutionary history, deployment is an individual lifetime. And our genomes are our reusable parameters. Unfortunately I haven't yet written any papers/posts really laying out this analogy, but it's pretty central to the way I think about AI, and I'm working on a bunch of related stuff as part of my PhD, so hopefully I'll have a more complete explanation soon.

2Steven Byrnes5y

Oh, OK, I see what you mean. Possibly related: my comment here.

[-]Richard_Ngo5y60

I've recently discovered waitwho.is, which collects all the online writing and talks of various tech-related public intellectuals. It seems like an important and previously-missing piece of infrastructure for intellectual progress online.

[-]Richard_Ngo6mo5-1

For a while now I've been thinking about the difference between "top-down" agents which pursue a single goal, and "bottom-up" agents which are built around compromises between many goals/subagents.

I've now decided that the frame of "centralized" vs "distributed" agents is a better way of capturing this idea, since there's no inherent "up" or "down" direction in coalitions. It's also more continuous.

Credit to @Scott Garrabrant, who something like this point to me a while back, in a way which I didn't grok at the time.

[-]Richard_Ngo5y50

Yudkowsky mainly wrote about recursive self-improvement from a perspective in which algorithms were the most important factors in AI progress - e.g. the brain in a box in a basement which redesigns its way to superintelligence.

Sometimes when explaining the argument, though, he switched to a perspective in which compute was the main consideration - e.g. when he talked about getting "a hyperexponential explosion out of Moore’s Law once the researchers are running on computers".

What does recursive self-improvement look like when you think that data might be t... (read more)

4Daniel Kokotajlo5y

Perhaps a data-limited intelligence explosion is analogous to what we humans do all the time when we teach ourselves something. Out of the vast sea of information on the internet, we go get some data, and study it, and then use that to make a better opinion about what data we need next, and then repeat until we are at the forefront of the world's knowledge. We start from scratch, with a vague understanding like "I should learn more economics, I don't even know what supply and demand are" and then we end up publishing a paper on auction theory or something idk. This is a recurisve self improvement loop in data quality, so to speak, rather than data quantity.

2Viliam5y

What counts as self-improvement in the scenario governed by data? You can grab the whole internet, including scihub and library genesis, and then maybe hack all "smart" appliances worldwide... and after that I guess you need to construct some machines that will perform experiments for you. But none of this improves the machine's "self". With algorithms, the idea is that the machine would replace its own algorithms by better ones, once it gets the ability to invent and evaluate algorithms. With hardware, the idea is that the machine would replace its own hardware by faster ones, once it gets the ability to design and produce hardware. But replacing your data with better data, that... we usually don't call self-improvement. Also, what kind of data are we talking about? Data about the real world, they have to come from the outside, by definition. (Unless they are data about physics that you can obtain by observing the physical properties of your own circuits, or something like that.) But there is also data in sense of precomputed cached results, like playing zillions of games of chess against yourself, and remembering which strategies were most successful. If this was the limiting factor... I guess it would be something like a bounded AIXI which hypothetically already has enough hardware to simulate a universe, it only need to make zillions of computations to find the one that is consistent with the observed data.

2Richard_Ngo5y

In the scenario governed by data, the part that counts as self-improvement is where the AI puts itself through a process of optimisation by stochastic gradient descent with respect to that data. You don't need that much hardware for data to be a bottleneck. For example, I think that there are plenty of economically valuable tasks that are easier to learn than StarCraft. But we get StarCraft AIs instead because games are the only task where we can generate arbitrarily large amounts of data.

[-]Richard_Ngo4y40

RL usually applies some discount rate, and also caps episodes at a certain length, so that an action taken at a given time isn't reinforced very much (or at all) for having much longer-term consequences.

How does this compare to evolution? At equilibrium, I think that a gene which increases the fitness of its bearers in N generations' time is just as strongly favored as a gene that increases the fitness of its bearers by the same amount straightaway. As long as it was already widespread at least N generations ago, they're basically the same thing, because c... (read more)

[-]Richard_Ngo4y*Ω34-2

A general principle: if we constrain two neural networks to communicate via natural language, we need some pressure towards ensuring they actually use language in the same sense as humans do, rather than (e.g.) steganographically encoding the information they really care about.

The most robust way to do this: pass the language via a human, who tries to actually understand the language, then does their best to rephrase it according to their own understanding.

What do you lose by doing this? Mainly: you can no longer send messages too complex for humans to und... (read more)

6johnswentworth4y

That doesn't actually solve the problem. The system could just encode the desired information in the semantics of some unrelated sentences - e.g. talk about pasta to indicate X = 0, or talk about rain to indicate X = 1.

2Gunnar_Zarncke4y

I expected you to bring up the Natural Abstraction Hypothesis here. Wouldn't the communication between the parties naturally use the same concepts?

4johnswentworth4y

Same concepts yes, but that does not necessarily imply that they're encoded in the same way as humans typically use language.

4RobertKirk4y

Another possible way to provide pressure towards using language in a human-sense way is some form of multi-tasking/multi-agent scenario, inspired by this paper: Multitasking Inhibits Semantic Drift. They show that if you pretrain multiple instructors and instruction executors to understand language in a human-like way (e.g. with supervised labels), and then during training mix the instructors and instruction executors, it makes it difficult to drift from the original semantics, as all the instructors and instruction executors would need to drift in the same direction; equivalently, any local change in semantics would be sub-optimal compared to using language in the semantically correct way. The examples in the paper are on quite toy problems, but I think in principle this could work.

4AprilSR4y

Not being able to send messages too complex for humans to understand seems to me like it’s plausibly a benefit for many of the cases where you’d want to do this.

1kave4y

steganographically?

2Richard_Ngo4y

Ooops, yes, ty.

[-]Richard_Ngo5y40

Greg Egan on universality:

I believe that humans have already crossed a threshold that, in a certain sense, puts us on an equal footing with any other being who has mastered abstract reasoning. There’s a notion in computing science of “Turing completeness”, which says that once a computer can perform a set of quite basic operations, it can be programmed to do absolutely any calculation that any other computer can do. Other computers might be faster, or have more memory, or have multiple processors running at the same time, but my 1988 A

... (read more)

[-]gwern5y150

Equivocation. "Who's 'we', flesh man?" Even granting the necessary millions or billions of years for a human to sit down and emulate a superintelligence step by step, it is still not the human who understands, but the Chinese room.

1[anonymous]5y

I've seen this quote before and always find it funny because when I read Greg Egan, I constantly find myself thinking there's no way I could've come up with the ideas he has even if you gave me months or years of thinking time.

3gwern5y

Yes, there's something to that, but you have to be careful if you want to use that as an objection. Maybe you wouldn't easily think of it, but that doesn't exclude the possibility of you doing it: you can come up with algorithms you can execute which would spit out Egan-like ideas, like 'emulate Egan's brain neuron by neuron'. (If nothing else, there's always the ol' dovetail-every-possible-Turing-machine hammer.) Most of these run into computational complexity problems, but that's the escape hatch Egan (and Scott Aaronson has made a similar argument) leaves himself by caveats like 'given enough patience, and a very large notebook'. Said patience might require billions of years, and the notebook might be the size of the Milky Way galaxy, but those are all finite numbers, so technically Egan is correct as far as that goes.

3[anonymous]5y

Yeah good point - given generous enough interpretation of the notebook my rejection doesn't hold. It's still hard for me to imagine that response feeling meaningful in the context but maybe I'm just failing to model others well here.

[-]Richard_Ngo4y30

It's frustrating how bad dath ilanis (as portrayed by Eliezer) are at understanding other civilisations. They seem to have all dramatically overfit to dath ilan.

To be clear, it's the type of error which is perfectly sensible for an individual to make, but strange for their whole civilisation to be making (by teaching individuals false beliefs about how tightly constraining their coordination principles are).

The in-universe explanation seems to be that they've lost this knowledge as a result of screening off the past. But that seems like a really predictabl... (read more)

7Vaniver4y

Tho, to be fair, losing points in universes you don't expect to happen in order to win points in universes you expect to happen seems like good decision theory. [I do have a standing wonder about how much of dath ilan is supposed to be 'the obvious equilbrium' vs. 'aesthetic preferences'; I would be pretty surprised if Eliezer thought there was only one fixed point of the relevant coordination functions, and so some of it must be 'aesthetics'.]

6Richard_Ngo4y

I don't think dath ilan would try to win points in likely universes by teaching children untrue things, which I claim is what they're doing. Also, it's not clear to me that this would even win them points, because when thinking about designing civilisation (or AGIs) you need to have accurate beliefs about this type of thing. (E.g. imagine dath ilani alignment researchers being like "here are all our principles for understanding intelligence" and then continually being surprised, like Keltham is, about how messy and fractally unprincipled some plausible outcomes are.)

[-]Richard_Ngo5y30

Half-formed musing: what's the relationship between being a nerd and trusting high-level abstractions? In some sense they seem to be the opposite of each other - nerds focus obsessively on a domain until they understand it deeply, not just at high levels of abstraction. But if I were to give a very brief summary of the rationalist community, it might be: nerds who take very high-level abstractions (such as moloch, optimisation power, the future of humanity) very seriously.

2adamShimi5y

It seems to me that the resolution to the apparent paradox is that nerds are interested in all the details of their domain, but the outcome that they tend to look for are high-level abstractions. Even in settings like fandoms, there is a big push towards massive theories that entails every little detail about the story. Though defining rationalist community as a sort of community of meta-nerds who apply this nerd approach to almost anything doesn't seem too off the mark.

2Dagon5y

I think you need to unpack "trust" and "take seriously" a little bit to make this assertion. I think nerds are generally (heh) more able to understand the lossiness of models, and to recognize that abstractions are more broadly applicable, but less powerful than specifics. I wouldn't say I trust or take seriously the idea of Moloch or the similarities between different optimization mechanisms. I do recognize that those models have a lot of explanatory and predictive power, especially as a head-start (aka "prior") on domains where I haven't done the work to understand the exceptions and specifics.

[-]Richard_Ngo6yΩ230

There's some possible world in which the following approach to interpretability works:

Put an AGI in a bunch of situations where it sometimes is incentivised to lie and sometimes is incentivised to tell the truth.
Train a lie detector which is given all its neural weights as input.
Then ask the AGI lots of questions about its plans.

One problem that this approach would face if we were using it to interpret a human is that the human might not consciously be aware of what their motivations are. For example, they may believe they are doing something for altr... (read more)

[-]Richard_Ngo6y*Ω130

I've heard people argue that "most" utility functions lead to agents with strong convergent instrumental goals. This obviously depends a lot on how you quantify over utility functions. Here's one intuition in the other direction. I don't expect this to be persuasive to most people who make the argument above (but I'd still be interested in hearing why not).

If a non-negligible percentage of an agent's actions are random, then to describe it as a utility-maximiser would require an incredibly complex utility function (becaus... (read more)

4TurnTrout6y

I'm not sure if you consider me to be making that argument, but here are my thoughts: I claim that most reward functions lead to agents with strong convergent instrumental goals. However, I share your intuition that (somehow) uniformly sampling utility functions over universe-histories might not lead to instrumental convergence. To understand instrumental convergence and power-seeking, consider how many reward functions we might specify automatically imply a causal mechanism for increasing reward. The structure of the reward function implies that more is better, and that there are mechanisms for repeatedly earning points (for example, by showing itself a high-scoring input). Since the reward function is "simple" (there's usually not a way to grade exact universe histories), these mechanisms work in many different situations and points in time. It's naturally incentivized to assure its own safety in order to best leverage these mechanisms for gaining reward. Therefore, we shouldn't be surprised to see a lot of these simple goals leading to the same kind of power-seeking behavior. What structure is implied by a reward function? * Additive/Markovian: while a utility function might be over an entire universe-history, reward is often additive over time steps. This is a strong constraint which I don't always expect to be true, but i think that among the goals with this structure, a greater proportion of them have power-seeking incentives. * Observation-based: while a utility function might be over an entire universe-history, the atom of the reward function is the observation. Perhaps the observation is an input to update a world model, over which we have tried to define a reward function. I think that most ways of doing this lead to power-seeking incentives. * Agent-centric: reward functions are defined with respect to what the agent can observe. Therefore, in partially observable environments, there is naturally a greater emphasis on the agent's vantage point in t

5Richard_Ngo6y

I've just put up a post which serves as a broader response to the ideas underpinning this type of argument.

4Richard_Ngo6y

I think this depends a lot on how you model the agent developing. If you start off with a highly intelligent agent which has the ability to make long-term plans, but doesn't yet have any goals, and then you train it on a random reward function - then yes, it probably will develop strong convergent instrumental goals. On the other hand, if you start off with a randomly initialised neural network, and then train it on a random reward function, then probably it will get stuck in a local optimum pretty quickly, and never learn to even conceptualise these things called "goals". I claim that when people think about reward functions, they think too much about the former case, and not enough about the latter. Because while it's true that we're eventually going to get highly intelligent agents which can make long-term plans, it's also important that we get to control what reward functions they're trained on up to that point. And so plausibly we can develop intelligent agents that, in some respects, are still stuck in "local optima" in the way they think about convergent instrumental goals - i.e. they're missing whatever cognitive functionality is required for being ambitious on a large scale.

2TurnTrout6y

Agreed – I should have clarified. I've been mostly discussing instrumental convergence with respect to optimal policies. The path through policy space is also important.

[-]Richard_Ngo6yΩ8130

Makes sense. For what it's worth, I'd also argue that thinking about optimal policies at all is misguided (e.g. what's the optimal policy for humans - the literal best arrangement of neurons we could possibly have for our reproductive fitness? Probably we'd be born knowing arbitrarily large amounts of information. But this is just not relevant to predicting or modifying our actual behaviour at all).

[-]TurnTrout3yΩ6120

(I now think that you were very right in saying "thinking about optimal policies at all is misguided", and I was very wrong to disagree. I've thought several times about this exchange. Not listening to you about this point was a serious error and made my work way less impactful. I do think that the power-seeking theorems say interesting things, but about eg internal utility functions over an internal planning ontology -- not about optimal policies for a reward function.)

2TurnTrout6y

I disagree. 1. We do in fact often train agents using algorithms which are proven to eventually converge to the optimal policy.[1] Even if we don't expect the trained agents to reach the optimal policy in the real world, we should still understand what behavior is like at optimum. If you think your proposal is not aligned at optimum but is aligned for realistic training paths, you should have a strong story for why. 2. Formal theorizing about instrumental convergence with respect to optimal behavior is strictly easier than theorizing about ϵ-optimal behavior, which I think is what you want for a more realistic treatment of instrumental convergence for real agents. Even if you want to think about sub-optimal policies, if you don't understand optimal policies... good luck! Therefore, we also have an instrumental (...) interest in studying the behavior at optimum. ---------------------------------------- 1. At least, the tabular algorithms are proven, but no one uses those for real stuff. I'm not sure what the results are for function approximators, but I think you get my point. ↩︎

2Richard_Ngo6y

1. I think it's more accurate to say that, because approximately none of the non-trivial theoretical results hold for function approximation, approximately none of our non-trivial agents are proven to eventually converge to the optimal policy. (Also, given the choice between an algorithm without convergence proofs that works in practice, and an algorithm with convergence proofs that doesn't work in practice, everyone will use the former). But we shouldn't pay any attention to optimal policies anyway, because the optimal policy in an environment anything like the real world is absurdly, impossibly complex, and requires infinite compute. 2. I think theorizing about ϵ-optimal behavior is more useful than theorizing about optimal behaviour by roughly ϵ, for roughly the same reasons. But in general, clearly I can understand things about suboptimal policies without understanding optimal policies. I know almost nothing about the optimal policy in StarCraft, but I can still make useful claims about AlphaStar (for example: it's not going to take over the world). Again, let's try cash this out. I give you a human - or, say, the emulation of a human, running in a simulation of the ancestral environment. Is this safe? How do you make it safer? What happens if you keep selecting for intelligence? I think that the theorising you talk about will be actively harmful for your ability to answer these questions.

2TurnTrout6y

I'm confused, because I don't disagree with any specific point you make - just the conclusion. Here's my attempt at a disagreement which feels analogous to me: My response in this "debate" is: if you start with a spherical cow and then consider which real world differences are important enough to model, you're better off than just saying "no one should think about spherical cows". I don't understand why you think that. If you can have a good understanding of instrumental convergence and power-seeking for optimal agents, then you can consider whether any of those same reasons apply for suboptimal humans. Considering power-seeking for optimal agents is a relaxed problem. Yes, ideally, we would instantly jump to the theory that formally describes power-seeking for suboptimal agents with realistic goals in all kinds of environments. But before you do that, a first step is understanding power-seeking in MDPs. Then, you can take formal insights from this first step and use them to update your pre-theoretic intuitions where appropriate.

9Richard_Ngo6y

Thanks for engaging despite the opacity of the disagreement. I'll try to make my position here much more explicit (and apologies if that makes it sound brusque). The fact that your model is a simplified abstract model is not sufficient to make it useful. Some abstract models are useful. Some are misleading and will cause people who spend time studying them to understand the underlying phenomenon less well than they did before. From my perspective, I haven't seen you give arguments that your models are in the former category not the latter. Presumably you think they are in fact useful abstractions - why? (A few examples of the latter: behaviourism, statistical learning theory, recapitulation theory, Gettier-style analysis of knowledge). My argument for why they're overall misleading: when I say that "the optimal policy in an environment anything like the real world is absurdly, impossibly complex, and requires infinite compute", or that safety researchers shouldn't think about AIXI, I'm not just saying that these are inaccurate models. I'm saying that they are modelling fundamentally different phenomena than the ones you're trying to apply them to. AIXI is not "intelligence", it is brute force search, which is a totally different thing that happens to look the same in the infinite limit. Optimal tabular policies are not skill at a task, they are a cheat sheet, but they happen to look similar in very simple cases. Probably the best example of what I'm complaining about is Ned Block trying to use Blockhead to draw conclusions about intelligence. I think almost everyone around here would roll their eyes hard at that. But then people turn around and use abstractions that are just as unmoored from reality as Blockhead, often in a very analogous way. (This is less a specific criticism of you, TurnTrout, and more a general criticism of the field). Forgive me a little poetic license. The analogy in my mind is that you were trying to model the cow as a sphere, but you didn

4TurnTrout6y

Thanks for elaborating this interesting critique. I agree we generally need to be more critical of our abstractions. Falsifying claims and "breaking" proposals is a classic element of AI alignment discourse and debate. Since we're talking about superintelligent agents, we can't predict exactly what a proposal would do. However, if I make a claim ("a superintelligent paperclip maximizer would keep us around because of gains from trade"), you can falsify this by showing that my claimed policy is dominated by another class of policies ("we would likely be comically resource-inefficient in comparison; GFT arguments don't model dynamics which allow killing other agents and appropriating their resources"). Even we can come up with this dominant policy class, so the posited superintelligence wouldn't miss it either. We don't know what the superintelligent policy will be, but we know what it won't be (see also Formalizing convergent instrumental goals). Even though I don't know how Gary Kasparov will open the game, I confidently predict that he won't let me checkmate him in two moves. Non-optimal power and instrumental convergence Instead of thinking about optimal policies, let's consider the performance of a given algorithm A. A(M,R) takes a rewardless MDP M and a reward function R as input, and outputs a policy. Definition. Let R be a continuous distribution over reward functions with CDF F. The average return achieved by algorithm A at state s and discount rate γ is ∫RVA(M,R)R(s,γ)dF(R). Instrumental convergence with respect to A's policies can be defined similarly ("what is the R-measure of a given trajectory under A?"). The theory I've laid out allows precise claims, which is a modest benefit to our understanding. Before, we just had intuitions about some vague concept called "instrumental convergence". Here's bad reasoning, which implies that the cow tears a hole in spacetime: The problem is that it's impractical to predict what a smarter agent will do, or wh

4Richard_Ngo6y

I'm afraid I'm mostly going to disengage here, since it seems more useful to spend the time writing up more general + constructive versions of my arguments, rather than critiquing a specific framework. If I were to sketch out the reasons I expect to be skeptical about this framework if I looked into it in more detail, it'd be something like: 1. Instrumental convergence isn't training-time behaviour, it's test-time behaviour. It isn't about increasing reward, it's about achieving goals (that the agent learned by being trained to increase reward). 2. The space of goals that agents might learn is very different from the space of reward functions. As a hypothetical, maybe it's the case that neural networks are just really good at producing deontological agents, and really bad at producing consequentialists. (E.g, if it's just really really difficult for gradient descent to get a proper planning module working). Then agents trained on almost all reward functions will learn to do well on them without developing convergent instrumental goals. (I expect you to respond that being deontological won't get you to optimality. But I would say that talking about "optimality" here ruins the abstraction, for reasons outlined in my previous comment).

2TurnTrout6y

I was actually going to respond, "that's a good point, but (IMO) a different concern than the one you initially raised". I see you making two main critiques. 1. (paraphrased) "A won't produce optimal policies for the specified reward function [even assuming alignment generalization off of the training distribution], so your model isn't useful" – I replied to this critique above. 2. "The space of goals that agents might learn is very different from the space of reward functions." I agree this is an important part of the story. I think the reasonable takeaway is "current theorems on instrumental convergence help us understand what superintelligent A won't do, assuming no reward-result gap. Since we can't assume alignment generalization, we should keep in mind how the inductive biases of gradient descent affect the eventual policy produced." I remain highly skeptical of the claim that applying this idealized theory of instrumental convergence worsens our ability to actually reason about it. ETA: I read some information you privately messaged me, and i see why you might see the above two points as a single concern.

2Pattern6y

Is the point that people try to use algorithms which they think will eventually converge to the optimal policy? (Assuming there is one.)

2TurnTrout6y

Something like that, yeah.

2DanielFilan5y

I object to the claim that agents that act randomly can be made "arbitrarily simple". Randomness is basically definitionally complicated!

2Richard_Ngo5y

Eh, this seems a bit nitpicky. It's arbitrarily simple given a call to a randomness oracle, which in practice we can approximate pretty easily. And it's "definitionally" easy to specify as well: "the function which, at each call, returns true with 50% likelihood and false otherwise."

2DanielFilan5y

If you get an 'external' randomness oracle, then you could define the utility function pretty simply in terms of the outputs of the oracle. If the agent has a pseudo-random number generator (PRNG) inside it, then I suppose I agree that you aren't going to be able to give it a utility function that has the standard set of convergent instrumental goals, and PRNGs can be pretty short. (Well, some search algorithms are probably shorter, but I bet they have higher Kt complexity, which is probably a better measure for agents)

2Vaniver6y

I'd take a different tack here, actually; I think this depends on what the input to the utility function is. If we're only allowed to look at 'atomic reality', or the raw actions the agent takes, then I think your analysis goes through, that we have a simple causal process generating the behavior but need a very complicated utility function to make a utility-maximizer that matches the behavior. But if we're allowed to decorate the atomic reality with notes like "this action was generated randomly", then we can have a utility function that's as simple as the generator, because it just counts up the presence of those notes. (It doesn't seem to me like this decorator is meaningfully more complicated than the thing that gave us "agents taking actions" as a data source, so I don't think I'm paying too much here.) This can lead to a massive explosion in the number of possible utility functions (because there's a tremendous number of possible decorators), but I think this matches the explosion that we got by considering agents that were the outputs of causal processes in the first place. That is, consider reasoning about python code that outputs actions in a simple game, where there are many more possible python programs than there are possible policies in the game.

2Richard_Ngo6y

So in general you can't have utility functions that are as simple as the generator, right? E.g. the generator could be deontological. In which case your utility function would be complicated. Or it could be random, or it could choose actions by alphabetical order, or... And so maybe you can have a little note for each of these. But now what it sounds like is: "I need my notes to be able to describe every possible cognitive algorithm that the agent could be running". Which seems very very complicated. I guess this is what you meant by the "tremendous number" of possible decorators. But if that's what you need to do to keep talking about "utility functions", then it just seems better to acknowledge that they're broken as an abstraction. E.g. in the case of python code, you wouldn't do anything analogous to this. You would just try to reason about all the possible python programs directly. Similarly, I want to reason about all the cognitive algorithms directly.

2Vaniver6y

That's right. I realized my grandparent comment is unclear here: This should have been "consequence-desirability-maximizer" or something, since the whole question is "does my utility function have to be defined in terms of consequences, or can it be defined in terms of arbitrary propositions?". If I want to make the deontologist-approximating Innocent-Bot, I have a terrible time if I have to specify the consequences that correspond to the bot being innocent and the consequences that don't, but if you let me say "Utility = 0 - badness of sins committed" then I've constructed a 'simple' deontologist. (At least, about as simple as the bot that says "take random actions that aren't sins", since both of them need to import the sins library.) In general, I think it makes sense to not allow this sort of elaboration of what we mean by utility functions, since the behavior we want to point to is the backwards assignment of desirability to actions based on the desirability of their expected consequences, rather than the expectation of any arbitrary property. --- Actually, I also realized something about your original comment which I don't think I had the first time around; if by "some reasonable percentage of an agent's actions are random" you mean something like "the agent does epsilon-exploration" or "the agent plays an optimal mixed strategy", then I think it doesn't at all require a complicated utility function to generate identical behavior. Like, in the rock-paper-scissors world, and with the simple function 'utility = number of wins', the expected utility maximizing move (against tough competition) is to throw randomly, and we won't falsify the simple 'utility = number of wins' hypothesis by observing random actions. Instead I read it as something like "some unreasonable percentage of an agent's actions are random", where the agent is performing some simple-to-calculate mixed strategy that is either suboptimal or only optimal by luck (when the optimal mixed strat

4Richard_Ngo6y

This is in fact the intended reading, sorry for ambiguity. Will edit. But note that there are probably very few situations where exploring via actual randomness is best; there will almost always be some type of exploration which is more favourable. So I don't think this helps. To be pedantic: we care about "consequence-desirability-maximisers" (or in Rohin's terminology, goal-directed agents) because they do backwards assignment. But I think the pedantry is important, because people substitute utility-maximisers for goal-directed agents, and then reason about those agents by thinking about utility functions, and that just seems incorrect. What do you mean by optimal here? The robot's observed behaviour will be optimal for some utility function, no matter how long you run it.

2Vaniver6y

Valid point. This also seems right. Like, my understanding of what's going on here is we have: * 'central' consequence-desirability-maximizers, where there's a simple utility function that they're trying to maximize according to the VNM axioms * 'general' consequence-desirability-maximizers, where there's a complicated utility function that they're trying to maximize, which is selected because it imitates some other behavior The first is a narrow class, and depending on how strict you are with 'maximize', quite possibly no physically real agents will fall into it. The second is a universal class, which instantiates the 'trivial claim' that everything is utility maximization. Put another way, the first is what happens if you hold utility fixed / keep utility simple, and then examine what behavior follows; the second is what happens if you hold behavior fixed / keep behavior simple, and then examine what utility follows. Distance from the first is what I mean by "the further a robot's behavior is from optimal"; I want to say that I should have said something like "VNM-optimal" but actually I think it needs to be closer to "simple utility VNM-optimal." I think you're basically right in calling out a bait-and-switch that sometimes happens, where anyone who wants to talk about the universality of expected utility maximization in the trivial 'general' sense can't get it to do any work, because it should all add up to normality, and in normality there's a meaningful distinction between people who sort of pursue fuzzy goals and ruthless utility maximizers.

Moderation Log

Richard Ngo's Shortform

6

Ω 3

1. The Derivatives Market: The Collateral Multiplier