"If we go extinct due to misaligned AI, at least nature will continue, right? ... right?"

This is a linkpost for https://aisafety.info/questions/NM1Y/If-we-go-extinct-due-to-misaligned-AI,-at-least-nature-will-continue,-right

[memetic status: stating directly despite it being a clear consequence of core AI risk knowledge because many people have "but nature will survive us" antibodies to other classes of doom and misapply them here.]

Unfortunately, no.^[1]

Technically, “Nature”, meaning the fundamental physical laws, will continue. However, people usually mean forests, oceans, fungi, bacteria, and generally biological life when they say “nature”, and those would not have much chance competing against a misaligned superintelligence for resources like sunlight and atoms, which are useful to both biological and artificial systems.

There’s a thought that comforts many people when they imagine humanity going extinct due to a nuclear catastrophe or runaway global warming: Once the mushroom clouds or CO2 levels have settled, nature will reclaim the cities. Maybe mankind in our hubris will have wounded Mother Earth and paid the price ourselves, but...

(See More – 354 more words)

2plex28m

This is correct. I'm not arguing about p(total human extinction|superintelligence), but p(nature survives|total human extinction), as this conditional probability I see people getting very wrong sometimes. It's not implausible to me that we survive due to decision theoretic reasons, this seems possible though not my default expectation (I mostly expect Decision theory does not imply we get nice things, unless we manually win a decent chunk more timelines than I expect). My confidence is in the claim "if AI wipes out humans, it will wipe out nature". I don't engage with counterarguments to a separate claim, as that is beyond the scope of this post and I don't have much to add over existing literature like the other posts you linked. Edit: Partly retracted, I see how the second to last paragraph makes a more overreaching claim, edited to clarify my position.

1jaan5h

i might be confused about this but “witnessing a super-early universe” seems to support “a typical universe moment is not generating observer moments for your reference class”. but, yeah, anthropics is very confusing, so i’m not confident in this.

plex10m20

By my models of anthropics, I think this goes through.

17quiet_NaN16h

I think an AI is slightly more likely to wipe out or capture humanity than it is to wipe out all life on the planet. While any true Scottsman ASI is so far above us humans as we are above ants and does not need to worry about any meatbags plotting its downfall, as we don't generally worry about ants, it is entirely possible that the first AI which has a serious shot at taking over the world is not quite at that level yet. Perhaps it is only as smart as von Neumann and a thousand times faster. To such an AI, the continued thriving of humans poses all sorts of x-risks. They might find out you are misaligned and coordinate to shut you down. More worrisome, they might summon another unaligned AI which you would have to battle or concede utility to later on, depending on your decision theory. Even if you still need some humans to dust your fans and manufacture your chips, suffering billions of humans to live in high tech societies you do not fully control seems like the kind of rookie mistake I would not expect a reasonably smart unaligned AI to make. By contrast, most of life on Earth might get snuffed out when the ASI gets around to building a Dyson sphere around the sun. A few simple life forms might even be spread throughout the light cone by an ASI who does not give a damn about biological contamination. The other reason I think the fate in store for humans might be worse than that for rodents is that alignment efforts might not only fail, but fail catastrophically. So instead of an AI which cares about paperclips, we get an AI which cares about humans, but in ways we really do not appreciate. But yeah, most forms of ASI which turn out for out bad for homo sapiens also turn out bad for most other species.

Monthly Roundup #18: May 2024

Zvi

As I note in the third section, I will be attending LessOnline at month’s end at Lighthaven in Berkeley. If that is your kind of event, then consider going, and buy your ticket today before prices go up.

This month’s edition was an opportunity to finish off some things that got left out of previous editions or where events have left many of the issues behind, including the question of TikTok.

Oh No

All of this has happened before. And all of this shall happen again.

Alex Tabarrok: I regret to inform you that the CDC is at it again.
Marc Johnson: We developed an assay for testing for H5N1 from wastewater over a year ago. (I wasn’t expecting it in milk, but I figured it was going to poke up somewhere.)
However,

...

(Continue Reading – 14190 more words)

Templarrr22m10

Europeans... vastly less rich than they could be.

POSIWID. Metric being optimized is not "having the most money". It is debatable if it should be, as one of the "poor Europeans" my personal opinion is that we're doing just fine.

Teaching CS During Take-Off

andrew carle

I stayed up too late collecting way-past-deadline papers and writing report cards. When I woke up at 6, this anxious email from one of my g11 Computer Science students was already in my Inbox.

Student: Hello Mr. Carle, I hope you've slept well; I haven't.

I've been seeing a lot of new media regarding how developed AI has become in software programming, most relevantly videos about NVIDIA's new artificial intelligence software developer, Devin.

Things like these are almost disheartening for me to see as I try (and struggle) to get better at coding and developing software. It feels like I'll never use the information that I learn in your class outside of high school because I can just ask an AI to write complex programs, and it will do it...

(See More – 541 more words)

Jiao Bu37m10

"[I]s a traditional education sequence the best way to prepare myself for [...?]"

This is hard to answer because in some ways the foundation of a broad education in all subjects is absolutely necessary. And some of them (math, for example), are a lot harder to patch in later if you are bad at them at say, 28.

However, the other side of this is once some foundation is laid and someone has some breadth and depth, the answer to the above question, with regards to nearly anything, is often (perhaps usually) "Absolutely Not."

So, for a 17 year old, Yes. &nbs... (read more)

Fund me please - I Work so Hard that my Feet start Bleeding and I Need to Infiltrate University

Johannes C. Mayer

Thanks to Taylor Smith for doing some copy-editing this.

In this article, I tell some anecdotes and present some evidence in the form of research artifacts about how easy it is for me to work hard when I have collaborators. If you are in a hurry I recommend skipping to the research artifact section.

Bleeding Feet and Dedication

During AI Safety Camp (AISC) 2024, I was working with somebody on how to use binary search to approximate a hull that would contain a set of points, only to knock a glass off of my table. It splintered into a thousand pieces all over my floor.

A normal person might stop and remove all the glass splinters. I just spent 10 seconds picking up some of the largest pieces and then decided...

(Continue Reading – 1763 more words)

Mo Putera40m10

These thoughts remind me of something Scott Alexander once wrote - that sometimes he hears someone say true but low status things - and his automatic thoughts are about how the person must be stupid to say something like that, and he has to consciously remind himself that what was said is actually true.

For anyone who's curious, this is what Scott said, in reference to him getting older – I remember it because I noticed the same in myself as I aged too:

I look back on myself now vs. ten years ago and notice I’ve become more cynical, more mellow, and more pro

... (read more)

4Mo Putera9h

Holden advised against this: Also relevant are the takeaways from Thomas Kwa's effectiveness as a conjunction of multipliers, in particular:

1Johannes C. Mayer14h

At the top of this document.

2Algon14h

Thanks!

D&D.Sci (Easy Mode): On The Construction Of Impossible Structures [Evaluation and Ruleset]

abstractapplic

41m

This is a followup to the D&D.Sci post I made last Friday; if you haven’t already read it, you should do so now before spoiling yourself.

Below is an explanation of the rules used to generate the dataset (my full generation code is available here, in case you’re curious about details I omitted), and their strategic implications.

Ruleset

Impossibility

Impossibility is entirely decided by who a given architect apprenticed under. Fictional impossiblists Stamatin and Johnson invariably produce impossibility-producing architects of the time; real-world impossiblists Penrose, Escher and Geisel always produce architects whose works just kind of look weird; the self-taught break Nature's laws 43% of the time.

Cost

Cost is entirely decided by materials. In particular, every structure created using Nightmares is more expensive than every structure without them.

Strategy

The five architects who would...

(See More – 227 more words)

The consistent guessing problem is easier than the halting problem

jessicata

This is a linkpost for https://unstableontology.com/2024/05/20/the-consistent-guessing-problem-is-easier-than-the-halting-problem/

The halting problem is the problem of taking as input a Turing machine M, returning true if it halts, false if it doesn't halt. This is known to be uncomputable. The consistent guessing problem (named by Scott Aaronson) is the problem of taking as input a Turing machine M (which either returns a Boolean or never halts), and returning true or false; if M ever returns true, the oracle's answer must be true, and likewise for false. This is also known to be uncomputable.

Scott Aaronson inquires as to whether the consistent guessing problem is strictly easier than the halting problem. This would mean there is no Turing machine that, when given access to a consistent guessing oracle, solves the halting problem, no matter which consistent guessing oracle...

(Continue Reading – 1047 more words)

SamEisenstat1h10

Note that Andy Drucker is not claiming to have discovered this; the paper you link is expository.

Since Drucker doesn't say this in the link, I'll mention that the objects you're discussing are conventionally know as PA degrees. The PA here stands for Peano arithmetic; a Turing degree solves the consistent guessing problem iff it computes some model of PA. This name may be a little misleading, in that PA isn't really special here. A Turing degree computes some model of PA iff it computes some model of ZFC, or more generally any $Σ_{1}$ theory capable o... (read more)

To get the best posts emailed to you, create an account! (2-3 posts per week, selected by the LessWrong moderation team.)

Stephen Fowler's Shortform

Stephen Fowler

Wei Dai1h20

It's also notable that the topic of OpenAI nondisparagement agreements was brought to Holden Karnofsky's attention in 2022, and he replied with "I don’t know whether OpenAI uses nondisparagement agreements; I haven’t signed one." (He could have asked his contacts inside OAI about it, or asked the EA board member to investigate. Or even set himself up earlier as someone OpenAI employees could whistleblow to on such issues.)

If the point was to buy a ticket to play the inside game, then it was played terribly and negative credit should be assigned on that bas... (read more)

1Stephen Fowler4h

This does not feel super cruxy as the the power incentive still remains.

4Joe_Collman6h

I think there's a decent case that such updating will indeed disincentivize making positive EV bets (in some cases, at least). In principle we'd want to update on the quality of all past decision-making. That would include both [made an explicit bet by taking some action] and [made an implicit bet through inaction]. With such an approach, decision-makers could be punished/rewarded with the symmetry required to avoid undesirable incentives (mostly). Even here it's hard, since there'd always need to be a [gain more influence] mechanism to balance the possibility of losing your influence. In practice, most of the implicit bets made through inaction go unnoticed - even where they're high-stakes (arguably especially when they're high-stakes: most counterfactual value lies in the actions that won't get done by someone else; you won't be punished for being late to the party when the party never happens). That leaves the explicit bets. To look like a good decision-maker the incentive is then to make low-variance explicit positive EV bets, and rely on the fact that most of the high-variance, high-EV opportunities you're not taking will go unnoticed. From my by-no-means-fully-informed perspective, the failure mode at OpenPhil in recent years seems not to be [too many explicit bets that don't turn out well], but rather [too many failures to make unclear bets, so that most EV is left on the table]. I don't see support for hits-based research. I don't see serious attempts to shape the incentive landscape to encourage sufficient exploration. It's not clear that things are structurally set up so anyone at OP has time to do such things well (my impression is that they don't have time, and that thinking about such things is no-one's job (?? am I wrong ??)). It's not obvious to me whether the OpenAI grant was a bad idea ex-ante. (though probably not something I'd have done) However, I think that another incentive towards middle-of-the-road, risk-averse grant-making is the last t

1starship0067h

Hmmm, can you point to where you think the grant shows this? I think the following paragraph from the grant seems to indicate otherwise:

On Privilege

shminux

The forum has been very much focused on AI safety for some time now, thought I'd post something different for a change. Privilege.

Here I define Privilege as an advantage over others that is invisible to the beholder. [EDIT: thanks to JenniferRM for pointing out that "beholder" is a wrong word.] This may not be the only definition, or the central definition, or not how you see it, but that's the definition I use for the purposes of this post. I also do not mean it in the culture-war sense as a way to undercut others as in "check your privilege". My point is that we all have some privileges [we are not aware of], and also that nearly each one has a flip side.

In some way this...

(See More – 332 more words)

14Viliam13h

What are the advantages of noticing all of this? * better model of the world; * not being an asshole, i.e. not assuming that other people could do just as well as you, if they only were not so fucking lazy; * realizing that your chances to achieve something may be better than you expected, because you have all these advantages over most potential competitors, so if you hesitated to do something because "there are so many people, many of them could do it much better than I could", the actual number of people who could do it may be much smaller than you have assumed, and most of them will be busy doing something else instead.

2shminux1h

I mostly meant your second point, just generally being kinder to others, but the other two are also well taken.

6localdeity3h

Also: * Knowing the importance of the advantages people have makes you better able to judge how well people are likely to do, which lets you make better decisions when e.g. investing in someone's company or deciding who to hire for an important role (or marry). * Also orients you towards figuring out the difficult-to-see advantages people must have (or must lack), given the level of success that they've achieved and their visible advantages. * If you're in a position to influence what advantages people end up with—for example, by affecting the genes your children get, or what education and training you arrange for them—then you can estimate how much each of those is worth and prioritize accordingly. One of the things you probably notice is that having some advantages tends to make other advantages more valuable. Certainly career-wise, several of those things are like, "If you're doing badly on this dimension, then you may be unable to work at all, or be limited to far less valuable roles". For example, if one person's crippling anxiety takes them from 'law firm partner making $1 million' to 'law analyst making $200k', and another person's crippling anxiety takes them from 'bank teller making $50k' to 'unemployed', then, well, from a utilitarian perspective, fixing one person's problems is worth a lot more than the other's. That is probably already acted upon today—the former law partner is more able to pay for therapy/whatever—but it could inform people who are deciding how to allocate scarce resources to young people, such as the student versions of the potential law partner and bank teller. (Of course, the people who originally wrote about "privilege" would probably disagree in the strongest possible terms with the conclusions of the above lines of reasoning.)

shminux1h42

Excellent point about the compounding, which is often multiplicative, not additive. Incidentally, multiplicative advantages result in a power law distribution of income/net worth, whereas additive advantages/disadvantages result in a normal distribution. But that is a separate topic, well explored in the literature.

The Natural Selection of Bad Vibes (Part 1)

Kevin Dorst

This is a linkpost for https://kevindorst.substack.com/p/the-natural-selection-of-bad-vibes

TLDR: Things seem bad. But chart-wielding optimists keep telling us that things are better than they’ve ever been. What gives? Hypothesis: the point of conversation is to solve problems, so public discourse will focus on the problems—making us all think that things are worse than they are. A computational model predicts both this dynamic, and that social media makes it worse.

Are things bad? Most people think so.

Over the last 25 years, satisfaction with how things are going in the US has tanked, while economic sentiment is as bad as it’s been since the Great Recession:

Meanwhile, majorities or pluralities in the US are pessimistic about the state of social norms, education, racial disparities, etc. And when asked about the wider world—even in the heady days of 2015—a only...

(Continue Reading – 1948 more words)

Kevin Dorst1h10

Agreed that people have lots of goals that don't fit in this model. It's definitely a simplified model. But I'd argue that ONE of (most) people's goals to solve problems; and I do think, broadly speaking, it is an important function (evolutionarily and currently) for conversation. So I still think this model gets at an interesting dynamic.

LESSWRONG
LW

Quick Takes

Popular Comments

Recent Discussion