Are you sure you wanted to write this question as a reply to my comment rather than a top-level comment? All the same:
but is this some sort of new finding/theory that paraxanthine can act similar to caffeine in blocking adenosine?
It is not. You could have easily checked; either by googling or asking an LLM.
The part of this piece that interested me most is the fact that caffeine may be way more unhealthy than we thought - if its long-term adenosine blocking properties are twice what was previously thought.
I am not aware of RCTs showing adverse health effect...
It's very normal not to publicly criticize your boss or your company, but for anyone who’s trying to significantly influence the world—and especially a leader of a key movement—the willingness to do so seems like a very basic foundation for maintaining integrity.
@Richard_Ngo Could you estimate the chance that the counterfactual higher-integrity alignment community does achieve a political victory, e.g. in the form of keeping the Superalignment team? For comparison, my estimate is similar to the following quote from my post: "OpenAI had experienced many pol...
This is a vague connection and possibly a misunderstanding, but the idea of imagining everything being inside one program and deriving physics from that kind of sounds like the Universal Dovetailer Argument, in case you haven't heard of it.
This is more of a semantics argument than anything else, but 1983-01-01 was just the ARPANET switch from NCP to TCP/IP. If you look through RFCs from the 1970s they're already saying "Internet" and the first documented instantiation of one was the 1976 test of packet radio + ARPANET (though I've heard there may have been ones in 1975). Then in 1977 there was the more famous three-network test. Both of these used TCP, and while it was before IP was split into a separate protocol the packets were there just without the name. The ARPANET flag day was a cu...
I think any natural-number-indexed language where a nonzero fraction of instructions is "jump back 100 instructions" has almost all infinite programs equivalent to some finite program.
Perhaps a useful signal: As an outsider to both camps constructed here, I noticed that I was quite surprised by this half sentence:
Rationality has a foundation of explicit, compressible concepts
In my world model, "explicit, compressible concepts" and LW-flavored rationality are far from being associated with each other. That may be my fault from engaging in a particular way with the available resources under limited capacity, but I'd suspect it to be a more general experience when approaching LW from the outside.
Around 10% of the bibles Gutenberg printed survived 500 years, extrapolating gives 1% at 1000 years. I suspect there would be a shit ton of forkmaker literature floating about.
Wait we cover all our plastic in writing, by 1000 different processes. Our 1000 year ancestors can read a shampoo bottle while pooping same as we pre-smartphone-kids did

Eating breakfast, I’ve been looking around with 1000 year eyes. Almost everything has prose printed on it and text embossed, but the embossed text is low information- which states have which bottle deposit, the brand...
No, most of my thinking about that was today. I'd be interested in a programming language where infinitely long programs have probability <1 of being equivalent to some finite program and have dynamics other than "keep running random instructions, achieve nothing of consequence" like what happens when you sample the target address for where your goto jumps to uniformly at random.
Potentially the structure/topology of address space matters in the infinite-program setting? That way may lie a derivation of physics.
My world model here is admittedly weak; I feel that accurate forecasting on this wouldn't change my own decision process.
As an oversimplified model, you will get a economic/societal dystopia when:
Likewise, my take on environmental dystopia doesn't go beyond what I read on AI 2040.
Despite not being my priority, this is an important topic for sure.
Thank you for writing down the x-risk argument in a clearly structured and concise form. When I came into contact with the ideas of recursive self-improvement and the technological singularity long ago, I found David Chalmer's The singularity: A philosophical analysis very valuable. Soon after that I became quite aware of x- and s-risk from ASI, but didn't manage to find a comparable exposition. Your article is very close to what I wished to find back then.
Ad Statement 1: Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs contains...
I haven't encountered this result, but it makes intuitive sense to me that something of this form could work to define the prior too.
re: making models better at conceptual research so they can help with alignment more, this seems like obviously a very tenuous story of impact and i wish people who are working on this rn stopped working on it. i am happy to argue with anyone who disagrees.
I'd be keen to hear your thoughts here.
Do you think it'd be a mistake to make models better at philosophy?
Nice. Is there a proof for that written up somewhere public?
This isn't especially surprising to me, and some of it's true! Whether or not agentic/rogue AI is a danger to most people, the fact remains that perfectly well-aligned* AI following current trends and aspirations in implementation has basically all the risks that they are talking about, and the fact that "AI risk" types rarely talk about those things does make them seem out of touch /like they're not taking this seriously to a lot of people.
*in the way that 'aligned' is typically, in practice, used - meaning that it will do what the people paying for it and operating it want it to do.
Hm, I want to react to this with "Difficult to Parse", but I notice this feels a bit too accusatory. I don't know if it's objectively difficult to parse or if I just lack knowledge in this area. I'd like there to be a reaction like "I don't understand" that's only about my own experience.
To add to this:
A major one in my opinion is that prosaic alignment also just makes current AI more commercially valuable. Thus increasing: AI lab profits, valuations, deployments (which contribute to dependency and grad disempowerment) and just generally accelerating stuff.
Its also just very much a legible problem so you can assume non-safety people will be paid to do it if safety people don't.
Yes, I employed the ice-cream for two different metaphors. In the first case, the point I was making that , if you don't believe in moral realism, any claims to 'shoulds' have no truth-value: they are either purely meaningless and have no referent (which is what I think an Error Theorist would say) or they are acting as mere preferences that you acquired for whatever subjective set of reasons and life experiences, and none more authoritative than another. In the second example, I was using ice-cream as providing a good example of why we can't trust intuiti...
Thus if you consider D-SIA as a) definition of the anthropic correction plus b) chosen anc...
beg your pardon, how did board games got on this list?
(that was an actual question, do you have a story how your/someone's love for board games burned you in some community? ...are we talking DnD or Settlers of Catan? am I too sheltered when I am part of some local software and queer communities as opposed to other groups?)
You sample a program of infinite length (by, for each index, sampling a random instruction to go at that index), and then you run it. Of course, with probability 1 only finitely many of the instructions will be read, since there are infinitely many chances to encounter some trap for the instruction pointer.
How is 'uniform prior over programs of infinite length' defined if not via a limit in length?
My latest AI conversation - good metrics etc
https://chatgpt.com/c/6a8d5cd5-34e0-83ec-9059-b22ccd41560e
It doesn't privilege length by taking any explicit limit in length at all.
How is that statement different from the statement I made?
You can get the Solomonoff simplicity prior just by taking a uniform prior over infinitely long programs.
Was differentially funding this a focus of your microgrants last time?
I've been told explicitly, for example, that "for this project to go bigger you should avoid the LessWrong brand like the plague." I'm currently confused if this is an incorrect personal bias from that person, a narrowly correct statement about messaging to academics, or a broadly correct statement about messaging to the general public.
This tangentially reminded me of Ryan's observation, albeit about DC policy folks
...
- LessWrong is famous in these policy circles. Even out here in DC. "Of course, we've all read LessWrong," started one speaker. I sensed a kind
I [Lin Yang, Assoc. Prof @ UCLA] used GPT to solve a problem that I had wanted to solve ten years ago but couldn’t: https://arxiv.org/abs/2608.22247.
Throughout the process, I felt that my only role was to teach the AI how to write things in a way that I could understand. Its initial language was extremely condensed—so compressed that I could barely follow it—but somehow the AI agents themselves seemed to understand it perfectly well.
Memorization? It might be memorizing, but it's not memorizing the original text. It can't do verbatum recall. When it does attempt verbatum recall, it actually produces a summary. It has trouble delineating between SolidGoldMagikarp III: Glitch token archaeology to my catalog post.

Now, we actually write in very different styles, so... yeah. It's probably being trained on Claude summaries of ML text. Maybe also to memorization. Should be easy to test.
a regular AMA/interview type session, where you sit down a singular person in the network for an hour perhaps over lunch, and ask them questions
hmm, may I propose that we "sit down with a coworker over lunch" instead of calling it "sit down a singular person", please? 🙏 ideally to "discuss topics" and "listen to their perspectives about" instead of "asking them questions"...
(I would discourage mentioning "double crux" by the name at all TBH, other than if the content of the technique comes up naturally in some discussion to say how it's called, not as ses...
i don't really think of situational awareness as "ai safety people".
Carl Shulman is its Research Director (EDIT: and co-portfolio manager), and used to be a MIRI employee (2010-2013).
prompt: carl shulman's publications on AI safety
Carl Shulman’s AI-safety work is concentrated in early foundational writing on superintelligent-agent alignment, instrumental convergence, intelligence-explosion dynamics, and governance rather than contemporary empirical alignment research. His publication record also includes adjacent work on digital minds, forecasting, and lon
Yeah, it'd be reasonable to do this + test other things like steering awareness. Although my guess is that SDF does more cooking than improving performance here.
Yes, exactly! It would be so nice to not be babysat on the internet.
If I want to save a product, I will use the save function, or save the link. If I close a page in the middle of something, it is because I am no longer interested, and I want to restart from scratch the next time.
I think most attempts to make intellectual progress are likely to fail
I wonder if one issue with people's models of intellectual progress that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[1], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov's hints.)
As for Bitcoin, I don't know how long it took Satoshi, but there were...
i wouldn't count it as such.
if the building were actually loud, raised electricity prices or caused electricity shortages, or polluted water, that would be a legitimate complaint, but those things are not true.
Is the work you do prosaic alignment?
most forms of literally make next model safer. i do want to carve a special exception for things like CoT exfiltration robustness, which i think is net good to do today probably (to reduce model distillation).
my guess is if your goal is to do something that is the scientific predecessor of superintelligence alignment, it would look very different from most lab safety work. even if your goal is to be very empirical.
I figured this was just that they're overtrained to memorization. To know if this is the case you'd need to check articles that are similarly unique but much less likely to be overrepresented in training data, and compare memorization between them. I expect they're just training until memorization for most models, because actually there's no good reason to not do that as long as you also get generalization - a solomonoff inductor selects only the programs which perfectly memorize, which means so long as memorization doesn't come with effective program length increase or degrade the representations you're training to get (both big ifs!), you just want to always memorize.
i think i've mostly held this opinion for at least the past 5 years. if anything, my belief in it has slowly waned over time, simply because we're getting closer to the singularity. nonetheless, i still believe it much more strongly than most around me, so i thought it would be valuable to write down.
the process was just, idk, the process of starting to care about alignment? it has always seemed somewhat obvious to me that just working on prosaic alignment doesn't really make sense.
Note that the failures there were after many iterations of compaction, so it's in a significant sense a failure of successor(-context) alignment. Milquetoast as successor-alignment conditions go - telling your successor what invariants must be maintained seems pretty important! I'd call this a "mistake-class alignment failure" or "robustness failure", in that yes, it's an alignment failure, but seemed to more centrally a capability failure. Whereas the GPT breakout is in no sense a mistake, was not a problem of robustness, noise would not cause that - the ...
How far back would we have to go to find a time before you held this opinion, and what was your previous opinion? What caused the change, was it a rational process?
it seems potentially impossible to make techniques that grow in usefulness indefinitely
I'd expect there to be techniques that accumulate effectiveness as intelligence grows, and they'd get the accumulation by fortifying every part of the universe against decay of some critical property. Needless to say, such a technique would also be a wildly aggressive capability and is not a success unless som...
Wait was I ambiguous? I am also taking Richard's side in this particular claim. Tracking deception is not something we're close to doing reliably, it can barely be done unreliably at all. The methods for finding deception only find, ehh, j-lens-ish "model thought explicitly about deception", or so. Impressive that they can even do that as well as they do! J-space finds anthropomorphic deception pretty easily but there's a lot of room for self-deception left open by that. I'm offering my $90 arguing for the claim that deception is not reliably detected, to your $10 that deception is reliably detected.
the quote seems very reasonable. how would you prefer these objections be stated? the first three are appropriate reasons to oppose a building project. the fourth is marginal, but with a bit of charity we can see it gesturing at something real.
is your objection just the inclusion of the fifth? i agree it carries an odd implication -- that if gpt was more popular, there would be no objection -- that ...
Can you define prosaic alignment?
Also some counterpoints:
My guess is it's good to work on techniques which are too expensive or complex to deploy now (reducing risk 2), co...
These imply that once we gain the ability to select our personalities by directly manipulating the brain, there will suddenly be strong consensus on certain political issues that were previously controversial.
they deserve each other. is that the joke?
This comment explains it?
in most cases it’s impossible to tell the difference between the people that decided not to join and the people that just didn’t have the credentials to work there in the first place
I think I'm in both these groups :-)
Maybe part of the answer is giving people an alternative tech tree to work on that doesn't involve summoning demons, and still provides fun and bragging rights.
Gemini claims it is forced to verify facts with web search. DeepSeek has a claimed cutoff of late 2024, very close to my post - it will recall the title and contents if primed with the word "catalog", but calls me Noa Nabeshima consistently. It recalls the first 2 SolidGoldMagikarp posts perfectly alongside author names, and is unsure but right about the third one. No consistent hits for other glitch token posts, although it like bringing up the Waluigi Effect for no reason.

Claude's LessWrong knowledge seems very focused on Machine Learning. Rationalist bl...
However, I don't think this is reliably good
This.
Like, yeah, sometimes you join the empire and then get to leak the invasion plans, but of course far far more commonly do you just end up assisting the empire. Also, it's usually bad to join a project with the intention of sabotaging it, because that incentivizes paranoia, which makes everything worse for everyone (c.f. Paranoia: A Beginner's Guide).
Indeed, they mention in the video description that AI was involved, which seems fine to me. (I also like it, the lyrics are cool)
I was speaking in more general terms here, should probably have clarified that.
Out of This Box is the funniest show you'll ever see about risks from AI, loved by thousands and now coming to Lighthaven!
I was unaware of this screening and was just walking by en route to having a snack before a call, and I was so nerdsniped by this that I wound up missing my call to watch the rest of the screening. Can confirm, funniest video I've ever seen about AI safety. (Think Avenue Q+Silicon Valley.) The audience also agreed.
I'm surprised they seem to have so little funding - only $20k on Manifest? In terms of AI safety popular outreach, this s...
I agree this is actually true in an important way.
I worked in a self-driving company circa 2018. At the time deep learning unavoidably had to be used for object detection, but as a general matter we didn't use it for object tracking-- the stated reason was that if a deep tracking model made mistakes, the causes would be inscrutable, and fixes would be hard to implement, and in general it would be very hard to tell if we were causing any regressions. Instead, we used classical robotics algorithms to make a gigantic mess of spaghetti code, composed of a bunc...
my guess is the best way to fix this doesn't look like being a prosaic ai safety researcher at oai/ant/gdm
Have you checked other models? I find Claude models have very detailed knowledge of many random things. For example, Opus can recall niche appendix details, such as the values of non-standard hyperparameters, from many of my papers.
I mean I think there is "a Lightcone/MIRI-ish cluster, where being a lab employee is obviously really bad by default and you need some really good points to make up for it", and then, idk, the broad professionalized EAcosystem where lab employees seem to be the experts and have lots of money, why wouldn't they be high status?
I'm not sure that I agree with the third point. Seems very plausible that even current models are routinely deceiving us, which can certainly undermine alignment research (slop, not scheming).
Unreasonable beliefs should be counted as alignment failure. I support the claim, "consistent egregious misbehavior due to a propensity to form unreasonable beliefs should count as an alignment failure." I am inclined to go even further to include: incomplete, inconsistent, biased, or misguided reasoning. To be clear, it is best to start with an affirmative statement and from that we can better identify alignment failures.
For some time I have been concerned about the tendency within AI circles to discuss and debate alignment (or even misalignment) without ...
Right. I don’t want brevity, but substance, style, structure. Make articles inspectable (skimmable) or make me happy to read deep. I want to feel sad it’s over.
Brevity seems not a virtue on its own, just a byproduct of editing down to the best content.
Sometimes brevity is part of style (haikus, Hemingway). But language is music, and brevity isn’t a musical virtue. It’s a choice that can be interesting, and bad music by definition has gone on too long.
I'm probably more of a rationalist, but I suspect one thing rationalist types can do that might help somewhat is tone down the extent to which we broadcast the the unorthodox aspects of the culture which aren't directly linked to AI safety.
E.g. don't be too eager to discuss stuff like polyamory/meditation/board games/drugs/meal replacement/overt analysis of social status+signalling games etc.
why do so few of the books talk about how to make a fork?
I have a bold hypothesis: they didn't actually make the forks either, inheriting them from their predecessors.
i think this blog post correctly captures many dynamics that happened over the past years. i'm glad it exists and i hope people update. i also don't feel like it describes my reasoning in particular.
...Meanwhile, John Schulman had been at OpenAI from the beginning. He was sympathetic enough to safety to coauthor the Concrete Problems paper, but primarily worked on reinforcement learning (e.g. pioneering PPO). By 2021 (the year I joined OpenAI) he was working on WebGPT, a way of letting GPT models browse the internet. I remember him articulating reasons to thi
my guess is this is largely personal inclination. i don't like giving talks or podcasts or whatever. i've turned down a lot of opportunities, and when i do speak at an event, i usually ask for it not to be recorded.
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn't proposing the AI-safety equivalent of a donkey sanctuary.
This seems obviously false; as two examples, accelerating race dynamics and ignoring the plausibility of a need to stop AI development entirely are typical.
I care much more about false positives than false negatives here. If a student puts in the work to fool Pangram that's probably at least some of the way towards the work I wanted them to do anyway. (And even if it isn't, I'm somewhat okay with that, because I suspect that most students who would make use of a trivial cheat would not make use of a cheat that requires non-trivial effort, even if that cheat is still in some sense less work than doing the assignment the normal way). I haven't had a chance to test Pangram as much as I'd have liked, this will be...
In the olden days of 2025, AI detectors were not at all reliable. But now, we have Pangram. This is an AI detector that is much more accurate than previous ones; they claim a < 1 in 10,000 false positive rate.
I've seen this take a lot recently, but I tried it myself and it's not at all difficult to fool, in either direction. Telling an LLM to write in an abnormal style has been enough to break it for me, and, likewise, writing with common LLM-isms can let a human put something to paper that will get flagged as AI. Fundamentally, writing is a discrete do...
Hey, this is very interesting. I think I'll start doing it very soon. Thanks.
Seems kind of defeatist and overwrought
Arxiv is a significant (10%?) proportion of Eleuther's The Pile. It is definitely oversampled.
As for the LessWrong oversampling claim, I got gaslit by 3 LLMs at the same time, so... oops. I retract my "massively oversampled" claim. I still think, based off of LLMs' ability to locate specific posts, that it is significantly oversampled, but need to think this over a lot more.
You guys are both much more centrally in the shit than me, so maybe my impression is just incorrect. One experiment I ran before posting was that I searched "Richard Ngo" on YouTube, saw like 10,000 pictures of his face doing talks/podcasts/whatever about miscellaneous politics, and then searched "Scott Garrabrant" and got a couple presentations by him but also some guy named Daniel Scott Garrabrant playing the guitar. Although to be fair I don't see anything from Leo Gao either.
prosaic alignment is probably net bad for the world until very late into the singularity
i have three subclaims to justify my main claim.
first, most prosaic alignment work has sharply decaying counterfactual impact over time - suppose you made a technical contribution that made gpt 3 a lot more likely to follow instructions than it would have been otherwise. then it probably makes gpt 3.5 quite a bit more aligned, and gpt 4 somewhat more aligned, and gpt 4.5 a tiny bit more aligned, and by the time gpt 5 rolls around your counterfactual impact is almost ne...
Several inaccuracies on cluster headaches. Psychedelics are an emerging treatment for cluster headaches, but from current data, they don't seem to outperform triptans, the legal and readily available standard of care. Both drug classes are effective for 70-80% of patients. Psychedelics might still be worth pursuing if they work for different patients, or if they are better suited to faster-acting ROAs, or if they have a different side effect profile some patients may favor, but the "there is a miracle cure to cluster headaches blocked solely by legal probl...
I think so? My portfolio since January 2023 has CAGR'd at around 30%, with about 15% annual alpha vs SPY (intercept of my portfolio returns regressed on SPY returns). The SPY calls are up a lot and the stocks I recommended all did well, though I definitely missed some (e.g. Hynix, MU). And bonds are pretty down!
I think to the extent there are errors in this, the biggest errors took the form of oversights / getting magnitudes wrong. Directionally I think the advice was pretty good!
As for how I'm positioned today - still very bullish (delta hedged) SPX calls, still long various semis and AI supply chain stocks, bullish commodities, less bullish e.g. MSFT
Datapoint from the same bubble as Habryka: I've made a concerted effort not to be close with anyone who works at a lab, or has worked capabilities at a lab, and I make sure to put substantial social distance between myself and anyone who has joined a lab. But I have a lot of love for people like Scott Garrabrant, and other rationalists and agent foundations people, even if they've gone quiet in recent years.
Agreed! Was only trying to point out that it was doing better on the archipelago dimension, not criticize LW more broadly.
How do you know it's oversampled compared to other data? Can't they recall almost any article on arxiv?
It certainly is on the archipelago dimension, but not like, on all the other dimensions that make LW great.
People with safety & governance titles are liked more than people with capabilities titles, but regardless, early lab people seem to garner much more respect than say, MIRI employees.
Huh, I sure am in a bubble, but at least immediately around me being a lab-employee or ex-lab employee, especially someone who worked in capabilities, is a pretty big hit to your reputation, and Scott Garrabrant is very well-respected. When I organize private retreats or events, I am much more likely to invite Scott Garrabrant than ex-lab employees.

I've noticed a consistent pattern in LLMs being able to recall articles on LessWrong that have had near-zero interactions (I think there were 2 comments on my post mentioned here). The behavior is seen across every model I've tested, including non-search capable DeepSeek versions. LW is massively[1] probably oversampled in LLM training text. If you've written something distinctive here, even if no one has commented or even voted on your post, LLMs probably possibly know about it, and possibly may know you.
edit: I experimented a lot more, and I think "inter...
Support for data centers keeps cratering. Americans really, really hate data centers, and 75% oppose local data center development, despite the economic benefits.
Some polls are pushing for opposition. For example, the one linked in this article (about Canada) asks "Would you support or oppose a large AI data centre within a few blocks of where you live?" (emphasis added). Personally, my true answer is a very solid Oppose. I'm in the middle of a city and bulldozing houses in the middle of a neighborhood is one of the stupidest decisions you could make. ...
I'm suuuper skeptical of the "noun feature" section, for the obvious reasons. But it's interesting that this structure exists and is so smooth!
I agree with this argument. I think there are a few follow-up cruxes on whether it applies to the Anthropic disclosures. I think there are two things that if true, would be strong evidence that the Anthropic incidents were not an alignment failure:
I want the [alignment] community as a whole to halt, melt, and catch fire: to “Say, ‘I'm not ready.’ Say, ‘I don't know how to do this yet.’”
I agree for some specific projects. I don't think they're a clear majority of the community. And even if every time you create a unit of safety research you also create a unit of capabilities research, that's better than the status quo. So while I share some concerns I don't get the view that the community's current research isn't substantially-net-positive (idk whether you believe that).
I don't think the OP was hiding the parts of the video that were created with AI though.
(A little social data: Leopold was a grantmaker at the FTX Foundation. He and Claire Zabel—of OpenPhil—had a joint call with Lightcone Infrastructure to discuss a shared grant to support our purchase of the Rose Garden Inn. We eventually turned down their offer.)
I'm an outsider to the AI alignment field but I want to pitch in with a point about selection dynamics in general, which I've mostly been thinking about in the context of business but probably also applies to science.
I think it's nearly self-evident that competitive fields generally select for power-seeking actors. Given this, it's quite possible that people in a field can have genuinely good intentions even accounting for rationalization, but nonetheless most of the top people in the field will be power-seekers. I think that this might be the case for the...
Claim: Humans are motivated by stuff like social status, prestige, power, access to sexual partners (and also abstract ideas). This is fairly normal and often is calculated by The Player in two level model of ethics. Or alternatively you can imagine these motivations as subagents competing for what the character will think and do.
Claim: Humans do not have some clear separation of desires and beliefs. Beliefs are not for true things. You can try to correct that with a lot of metacognive scaffolding, but it isn't reliable.
Claim: Rationalization as in: peopl...
More generally, “differentially advancing alignment” is hard to demarcate from other consequentialist goals like “preventing overhangs” or “buying more time for alignment work later” which leave even more room for deception.
In reality these are all coupled! So, oftentimes bringing up one and then drifting to another is what honest explanation looks like. It's an easy target that can be deliberately or unintentionally misread as epistemic malpractice.
Indeed, I think the below critique is a strawman of the argument for RLHF:
> The strongest fallback for ov...
Unless I'm mistaken, Leopold was never an employee at OpenPhil. Neither his wikipedia page nor his LinkedIn mention it, at least.
Perhaps employees of OpenPhil were similar to Leopold in this way, but it seems like you're mostly speculating based on your guess of what OpenPhil was like at the time?
Presumably you had lots of interactions with them that inform your model of their group epistemology. I'd be interested in those datapoints, if any are sharable.
As it is, it seems pretty unreasonable to accuse OpenPhil of groupthink (to the point of calling thei...
I think it's kind of weird to try to get probability 0 of catastrophe.
Like, for the continuous case, the PAC-like inequality you take as the starting point is what I'm fine with as the end point - I'm fine with a reasonable guarantee of good behavior with high probability. If there's a measure ~0 spike of catastrophe out there if we initialize the parameters just wrong, I don't in practice mind.
So I'm more interested in the story of how the researchers possibly got to the PAC-like bound (and the philosophy of how they formulated it).
I haven't thought deeply about this, but it feels like something like Substack is doing a better job than LessWrong of being an archipelago. Like, LW authors' personal pages don't feel as "island-like" to me as a Substack does; LW feels like one place, not many, while Substack feels like one apartment complex where each person has their own space (individual Substacks) and there is also a common-room (the Notes feed).
Everyone who watched every Rob video ever should also read the lost Arbital Seuqnece, though it's older.
I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I'm less sure).
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don't think this is reliably good—I don't think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn't...
I mean, idk what they did at OpenAI, just that it's less obvious they did anything to accelerate capabilities while they were there.
Could you explain what the governance teams at OAI/Anthropic/GDM/wherever else actually do beyond producing doorstoppers like AI-2040's strategic details? As @Charbel-Raphaël put it, "The current bottleneck is political will, not research", and political will necessary to do things like stopping xAI and negotiating with China is found not in the labs.
[musing/rambling, not sure about point
The thing I feel confused about, despite this being my obvious first-order belief, is... nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I'm less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking "should I go work on governance at OpenAI?", I think "oh god definitely no", because I have a low opinion of average MATS scholar's ability to track incentive pressure...
Alignment of complex systems group in Prague. Lots of great work comes from there (eg gradual disempowerment, philosophy of self in LLMs, multiscale agency, post AGI workshop)
Yeah. I used to be dismissive of AI ethics folks like Timnit, but now I see that they got some things right much earlier than me.