Not a full solution, but one obvious thing that helps on the margin with improving the external image of "Wentsworth's club" is putting serious effort into making clear, accessible, and visible explanations of work being done within the group, why the group believes certain alternative approaches are [infeasible, unrealistic, flawed, etc.], background assumptions of the work, ideas that the group generally accepts vs. ideas that are controversial within the group, etc.
I have a low sample size, but of the people I've talked to that are doing work that seems to me like streetlighting, their responses to me bringing up something like Agent Foundations are usually less "ah yes Agent Foundations, that ridiculous cabal of groupthinkers" and more "what's Agent Foundations?" followed by what I assume is genuine curiosity and eagerness to learn more about AF ideas. (Agent Foundations is just an example; I'm aware that the alignment landscape is far more vast and complex than a naive prosaic vs. AF dichotomy).
This is slightly biased by the fact that I'm mostly talking to people who are relatively new to the field of alignment (as I am). But I think it's telling of the fact that a lot more pedagogical work could be done to improve the image of the parts of the field tackling some of the harder problems in alignment. My own experience also supports this idea. Coming from the EA/80k-hours side, my image of the field of alignment for the first 6 months was basically: "go get good at ML engineering and work for Anthropic on interpretability or scalable oversight or something along those lines", and I only got out of this epistemic state through a genuine effort to seek out other perspectives on the alignment problem. In hindsight, this seems problematic that I had to do this.
Having more people working on pedagogy also works on lessening another epistemic trap: the loop that goes something like "streetlighting research is easy and publishable and aligns with lab incentives" -> "more funding and jobs here" -> "more people whose careers or job prospects depend on believing optimistic alignment assumptions, and more visible research published" -> "more funding and jobs towards streetlighting because of visibility and epistemic cascades". It seems extremely difficult to get a job working on the hard problems of alignment, and this unfortunately doesn't seem easily solvable: the hard problems of alignment are, well, hard, and it takes a special talent to work on them. It seems less hard (but still some hard) to improve the pedagogy of the field, and thus this could be a more accessible option to people who want to contribute to the field but don't want to go work in one of the major labs. I don't know how many full-time positions like this could be created, but it seems like it could lessen the impact of the career-incentives loop above at least a little bit.
Lastly, not everyone can be as prolific as you Steve :), so this could lessen the burden of making their work accessible on some researchers.
I think a problem here might partially be that 95%+ of resources seems to be going to education about the initial view you described? E.g Bluedot, MATS and a lot of the introductions to AI Safety nowadays are focused on the things you described.
There are various reasons for this happening and I want there to be more diversity. I don't think it is necessarily fair to say that the way to solve this is through theory focused researchers to explain their ideas in more detail although that might be part of it.
Who sets the foundational incentives? Who is it pushing money into various parts of AI Safety?
I think funders need to get a less impoverished model of philosophy of science than they currently have and shift their distribution of bets.
The point in my mind is then pointing more at convincing funders at this level rather than it is to create better communication of any specific AF approach at least if you want to tackle what you're bringing up?
I think we agree on a fair amount here. The problem of improving the field's epistemics is complex and multifaceted, and I think both interventions (and more) should be tried. Indeed I think if multiple points of attack are not pursued, progress will be slow because there are interlocking problems and annoying feedback loops like the one I mentioned above.
Also to clarify, I'm not sure if "communicating a specific AF approach" is capturing all of my idea here. I'm more envisioning a general pedagogical/epistemic effort, covering additional things like open questions/disagreements vs. settled ideas/agreements within the group, explanations of why other approaches seem to not work, addressing of counterarguments, organization of material so that it's easily accessible/findable, synthesization of scattered arguments on LW that when put together form coherent pictures, analysis of background assumptions, research philosophy, reading lists, etc.
I know that all this work is happening on a small-scale right now, but I think significantly more could be done. Especially if, like you mentioned, more funding was directed towards efforts like this. (A feedback loop: less funding for work like this -> less work like this -> less informed funders -> less funding for work like this).
I do think we agree on a bunch and I didn't mean to imply that you didn't think the systemic funding part was important.
I think the signal I got in the back of my head was something similar to that when I hear people talk about how in order to reduce climate change we should all live more frugally when industrial operations account for like twice the CO2. Yes reducing personal emissions is good but developing bio-friendly concrete is a lot better.
I think one of the problems I still see with the "Communicate it better" (apologies for strawman summary!) solution is that a lot of weird theoretical research is somewhat ineffable? There's just a large inferential gap a lot of the time and so it is unclear exactly what is happening there and part of making a science paradigmatic is what you're pointing at which is this sort of settling into a stable paradigm where we have agreements and disagreements clearly mapped out. In my personal experience it is just too wild of a west out here in theory land to be able to do so easily, the dimensionality of the space is just really large.
I would love there to be a systematic distillation that makes sense of these issues and I also think that at least half of the community would disagree with such a general distillation?
What are the most important primitives? What types of questions are the main questions we should ask in doing this? What is the method in which we make progress in agent foundations? There are a couple of different answers for each of these questions afaik.
I would like to say that we should just let it wild west for a bit and be okay with not everything being so legible, that seems to be the way that theoretical fields in the past have made good progress for in the systemisation of what is allowed you also lose part of the creativity and flexibility needed to tackle deep problems?
I've recently been reading Theory and Reality by Peter Godfrey Smith and Against Method by Paul Feyerabend. I really like the first book as it is a great history of philosophy of science and I think the second book makes some very interesting points about work on the frontier of science that we do not understand!
(I think they point at why doing this is such a hard thing which might be interesting for you.)
I'm not disagreeing that it would be great to have it, I'm just not certain we are ready for it yet?
Hey, sorry for the late reply; I was on vacation. Thanks for the book recommendations, I definitely am looking for arguments to stress test this point I'm making so I'll definitely check those out!
To reply to the rest of your comment, I think I'm maybe hoping for more marginal progress than you seem to think I am?
There's just a large inferential gap a lot of the time and so it is unclear exactly what is happening there and part of making a science paradigmatic is what you're pointing at which is this sort of settling into a stable paradigm where we have agreements and disagreements clearly mapped out. In my personal experience it is just too wild of a west out here in theory land to be able to do so easily, the dimensionality of the space is just really large.
I would love there to be a systematic distillation that makes sense of these issues and I also think that at least half of the community would disagree with such a general distillation?
What are the most important primitives? What types of questions are the main questions we should ask in doing this? What is the method in which we make progress in agent foundations? There are a couple of different answers for each of these questions afaik.
Like here, I agree with a lot of the issues you bring up with the more ambitious version of improving the field's epistemics (that of making the field paradigmatic rather than pre-paradigmatic), but mostly what I am hoping for is just slightly more agreement, slightly better focus, etc., which I think is achievable. I still would want effort to be put into the more ambitious version as well; I think there are reasons for hope on this front (LW is special compared to other scientific dialogues, a lot of previous effort on this front has been done by lead researchers who are otherwise quite busy, I have not seen the kind of efforts that have looked systemic and thorough enough to me to be successful, and there may be some better ways of making progress on this that we haven't thought of yet). However, I would need to be able to respond to the arguments that you are making, the arguments that Steve made in his OP, and the arguments in those 2 books you recommend to properly defend this more ambitious effort.
I think the signal I got in the back of my head was something similar to that when I hear people talk about how in order to reduce climate change we should all live more frugally when industrial operations account for like twice the CO2. Yes reducing personal emissions is good but developing bio-friendly concrete is a lot better.
I think this argument bites harder if I'm hoping for more marginal progress, however. I'm not sure how to respond to it? It seems like we have different intuitions for how valuable this kind of work is. I guess I could point to Byrnes' recent pedagogical work and suggest that it seems to be somewhat effective? If you have any idea how to unpack this crux (if it is one, I'm assuming it is here but correct me if I'm wrong), let me know.
On the other hand, I would love to know the ideas you have for improving the funding landscape that don't route through more convincing arguments. I can't seem to think of many, but I'm probably not thinking creatively enough.
for in the systemisation of what is allowed you also lose part of the creativity and flexibility needed to tackle deep problems?
I agree that this is a concern, although less so if we are just aiming for marginal progress. I'll have to think about this more. I'm overall not too concerned about it, but that's just an intuition so far.
I suppose I'm being somewhat annoying and provocative and taking the more extreme versions of the points I could make in order to try to point out a different way of viewing the problem.
I completely agree with you that communications work is great and I think that anytime that Steven Byrnes or John Wentworth or any of the existing Agent Foundations people explain their research in more easily understandable ways that helps.
It's a good point and it should be done. Natural Abstractions: Key Claims, Theorems, and Critiques is one of my favourite posts on this site. It is a distillation post of the natural abstractions agenda, it's really good.
I think the core point I wanted to point at was something like that the research process is inherently coupled to the communication process. Communicating something clearly often happens only after you understand it deeply yourself, so in trying to communicate something weird you often make most of the progress on the abstractions. To give an example, I've been trying different angles on communicating a sort of very specific program that combines complex systems sciences with agent foundations. The pedagogy and making people understand claims is like the absolute hardest part of it. There is just a really large inferential gap if I want people to follow me on this journey.
After doing it long enough I recently stumbled on a neat little proof that bounds the size of collective agents (think an AI Swarm) to their communication speed and network topology which is an example of how this area of research is interesting. My plan is now to just communicate that example and through that point at the usefulness of the agenda as people won't listen if I start explaining graph signal processing to them as they think it is academic and with no relevance what so ever to AI Safety.
I suppose the point of disagreement is how you get to a point of being able to communicate something clearly rather than whether we should spend more resources and effort on communication?
There's also a second point where you try to explain something clearly but in doing that you lose some of the underlying nuance? Schrödinger's cat was a way to make people understand quantum mechanics on a basic level but they missed a lot of the math that is actually required to fully understand the problem? How do we know when a concept can be simplified in that way and still contain the true underlying problem?
Like here, I agree with a lot of the issues you bring up with the more ambitious version of improving the field's epistemics (that of making the field paradigmatic rather than pre-paradigmatic), but mostly what I am hoping for is just slightly more agreement, slightly better focus, etc., which I think is achievable. I still would want effort to be put into the more ambitious version as well; I think there are reasons for hope on this front (LW is special compared to other scientific dialogues, a lot of previous effort on this front has been done by lead researchers who are otherwise quite busy, I have not seen the kind of efforts that have looked systemic and thorough enough to me to be successful, and there may be some better ways of making progress on this that we haven't thought of yet).
I guess this is the process of the field becoming more paradigmatic for me? The concepts are simpler to understand, the problems are simpler and so it is easier to direct focus and to do science on it?
I think this argument bites harder if I'm hoping for more marginal progress, however. I'm not sure how to respond to it? It seems like we have different intuitions for how valuable this kind of work is. I guess I could point to Byrnes' recent pedagogical work and suggest that it seems to be somewhat effective? If you have any idea how to unpack this crux (if it is one, I'm assuming it is here but correct me if I'm wrong), let me know.
Just to make it clear, this was more of the vibe I was feeling. I'm not sure I understand what you're saying here with it biting harder if you're hoping for more marginal progress? Maybe my answer above answered some of these things?
On the other hand, I would love to know the ideas you have for improving the funding landscape that don't route through more convincing arguments. I can't seem to think of many, but I'm probably not thinking creatively enough.
Maybe through hosting programs and seminars where scientists get a lot of time to look at the problem and try to find new ways of examining it and communicating it? I see it a bit more as a puzzle than a sort of linear process, it is not like Hubinger's risks from inner optimisation did not introduce new concepts in order to explain existing problems within the field. A good distillation introduces new ways of viewing the world and so you need new ways of viewing the world to write a really good distillation.
For me this is one of the core points of good scientific apprenticeship, learning how to communicate something complicated easily. I would not say that I'm that good at it myself yet but I'm better than I was a year ago.
So I completely agree with your program and thinking here, I think it just turns out to have lots of annoying strange conundrums in the details?
I like this model as something to keep in mind, but don't think that string theory is an example unless I'm misunderstanding. I'm someone who worked on related math but doesn't know the physics super well, so might have gaps compared to a string theorist. But I don't think I've ever observed doubt that there was a meritocracy, or that it was opaque to non-string-theorists. My sense is that people in math and physics were excited about exactly what you say (that it in principle could be a field theory that allows for gravity) and it did get a lot of funding/ interest because of this (maybe at first it took some time to percolate as an exciting subject). So by the time you say it was "breached" it was very non-opaque, and I think acknowledged as a meritocracy (i.e. "the coolest high-energy physicists work in it"). In fact a phenomenon that was happening extremely strongly here, that happens whenever a kind of science is recognized as "sexy", is that it started accumulating less-meritocratic/less-careful stuff. And also as you point out, it developed pointless politization and "edgy" negative press by the likes of Woit.
But I think the thing that made people less excited was driven precisely by the "meritocratically selected" high-energy physicists who worked in it (or have the context to understand it) and got disappointed. Not by a few edgy blogs. I think that the thing that people found disappointing is that it failed to make contact with reality - exactly as you say, that it didn't generate a visible artifact, and moreover hadn't made progress that could lead to such an artefact for a long time.
My sense is that there are two main sources of disappointment wrt its ability to make contact with reality. One is that the energies associated with it are so high, so direct experiments are probably impossible, and the second is that there aren't many good theories that have the kind of non-perturbative structure (both UV and IR) that our universe has. There was ADS-CFT but not much beyond that.
So my sense is that the belief among the people within string theory is that yes, probably quantum gravity, or at least something in the direction of quantum gravity requiers string theory. But the most honest of these people think that ambitious "guess the theory of everything" (via the right Calabi-Yau and working out all the IR/UV issues) version of string theory is probably not the directions most worth pushing now - rather the interesting directions are making more sense of more localized questions. Make small progress in interesting theory QFTs, understand more about the boundary between GR and QFT (especially black holes, dark energy, big bang etc.) in our universe, and this is more likely to "raise the sea of knowledge" than immediately going for a general string theory.
Also another thing that I think happened is that to some extent string theory is now just one of many ways of getting QFTs now that people study. It's just not trumpeted as a fancy thing, but is a standard-ish tool in a high-energy physicist's arsenal.
BTW my colleague Lauren Greenspan (who would also have a lot more context here) has a related paper, where she studies interactions between string theorists and experimentalists in a history/ philosophy-of-science framework: https://arxiv.org/pdf/2205.05159
Are you saying this as criticism of the essay or in its favor? (This is not entirely a rhetorical question, unfortunately)
My gut instinct is that the whole thing reeks of a false dilemma on top of a strawman and the author isn't really acknowledging how strong of a prior scientific consensus is. Are we supposed to be convinced of this idea because if it were true, that would explain a lot?
I genuinely don't understand how this has so many upvotes (106 at the moment of writing the comment) and I'm not happy with what it would say about this community.
I'm reserving judgment, but man am I confused!
Here's why I upvoted:
The future partially hangs on this issue. The alignment community is saying "advanced AI is extremely dangerous". We claim to be a meritocracy of experts, with the more expert among us typicallly expressing levels of concern far above those of people outside this group.
If the public and politicians take us to be in the first category, a self-dealing cabal enmeshed in groupthink, they will ignore our concerns and press forward. And we all may die. If they take us to be the actual experts on the issue, there's a much larger chance they'll actually find ways to slow down progress at least somewhat, improving our odds of success a little, or perhaps even a lot.
Therefore, understanding the tension might help us figure out how to be interpreted as an actual externally opaque meritocracy.
Having said that explicitly and convinced myself, I'm upgrading from an upvote to a big upvote.
Okay, now I am extremely surprised.
I genuinely and sincerely did not believe that the previous statement was said in a positive light.
Are you getting behind an unfalsifiable claim? If the string theory example were false, how would you be able to find out?
This is very important to answer, unless you're renouncing yourselves as a meritocracy of critical thinkers (slightly ambiguous statement. I don't mean to say that Seth represents the community).
Were you just reading the essay, thinking about AI alignement the whole time, not particularly caring about the veracity of the specific examples?
I think Seth understood very well that you were criticizing my OP, and he changed his vote in the opposite direction to flaunt how strongly he was disagreeing with you.
(Just facilitating, not taking sides.)
(But I’m interested if you could spell out what you’re referring to by “unfalsifiable claim” in that comment.)
The way I read Seth's initial comment is that arguing that (for example) string theorists aren't an externally-opaque meritocracy, only illustrates the point that they are. That would take the claim to be unfalsifiable. I.e you'd reject concrete counter-evidence because it actually confirms the claim using this abstract argument.
Let me say outright: That's crazy. I don't think you agree with that. And if I understand Seth's reply just below, I now believe Seth also disagrees with that (phew!).
I'd still appreciate an answer to "if the string theory example were false, how would you be able to find out" by Seth. Just as a sanity check, for me. It's a routine question in SE and it shouldn't be hard to answer. I reiterate that isn't about whether string theory is true or false or whatever (which is what Seth now addressed a little further below), but about the experts dynamic illustrated (hypothesized?) here.
Disclaimer: I think the string theory example is wrong and have what I believe to be reasonable counter-evidence to back that up (again, not evidence that String Theory is wrong!).
I honestly haven't read the analytic philosophy example because the string theory one already raised quite a bit of red flags for me. I think I'm slightly less worried, but the trees and forest analogy down below is a new red flag! A forest of rotten trees isn't one where I want to be. You wouldn't get away with rotten trees in academia. I'd prefer if LW plainly rejected rotten trees and frowned upon possibly/possibly not rotten trees.
Disclaimer 2: I'm currently operating on gut feelings, because the whole thing rubbed me off. I'll probably take it slow soonish.
I didn't take it as criticism of the OP but just confusion. In the course of trying to clear that up, I realized I think this post is more important than I originally gave it credit for. I read it quickly and then had to move on yesterday.
This post was a little more confusing than your standard very clear writing, but I think it needed to be in order to give the experience of the outside perspective on fields that could be either genuine expert meritocracies or self-serving groupthink, before polluting that mental exercise with our probably-strong feelings about the field of alignment. So some extra digestion was needed. I'm not surprised that Pawn and perhaps the others who just addressed the object-level examples were confused about the larger point.
The essay says it's not about the specific examples, so yes I was not focused on that. The specific examples illustrate how the rest of the world must look at the field of alignment.
My comment was a somewhat restrained way of expressing frustration that most of the comments were about the specific examples, and thus missing the forest for the trees.
I worry that tree-focus is rampant in this community and another factor that could help get us all killed.
I don't know how I'd figure out whether string theory is or isn't a good direction in physics. This post doesn't really propose to answer the question of how to tell the difference from the outside or make it legible from the inside. But it raises an important question which probably most of us haven't thought much about, and that's the first step to getting important answers.
Said Achmiz has an old comment (which I can go dig up iff that would be helpful), saying something to the effect of "if the case for a phenomenon rests on some examples the author has provided, and you refute the examples, then until further examples are provided there isn't actually a case for the phenomenon." I don't think Pawn is claiming to have refuted the examples, but I suspect a similar instinct lies behind their comment: if the case for a phenomenon rests on the examples, then the examples matter a lot actually, and saying that criticism of the examples is "missing the forest for the trees" can feel like choosing the bottom line without reference to the actual quality of the arguments being given.
I think I actually disagree with Said here, and am inclined to think that even examples where, if you dig into them in detail, they don't quite fit as examples of the claimed phenomenon, can still serve to illustrate the phenomenon well enough to argue for its existence and importance. But I'm not exactly sure how to disagree; Said's point also feels at least partly right. Not exactly sure how to reconcile.
Examples can be explicitly made-up and illustrative, e.g. an infinite perfect crystal in physics class. The point isn’t to prove that the hypothesis is true, but rather to explain what the hypothesis is. (In many cases that’s 99% of the work. Again, think of physics courses; >99% of the time is spent trying to get people to understand what the laws of physics are and what they imply, and <1% of the time is spent doing the Bayesian thing where there’s multiple candidates for the laws of physics and you figure out which one is true by comparing their predictions about some piece of data.)
Somewhere between explicitly-unrealistic examples like infinite perfect crystals, and completely-realistic examples, we have stylized real-world examples, where the details don’t all stand up to scrutiny, but it’s still good enough to help point to a hypothesis. Like, maybe it’s a thing that could have happened, or it’s one part of a more complicated story, or whatever.
(But I’m certainly interested to know whether my OP examples are misleading, and would change the post if I came to believe that.)
the whole thing reeks of a false dilemma on top of a strawman
If you want to spell out what you mean by that, I’d be interested.
I genuinely don't understand how this has so many upvotes
For the record, I sure don’t think this post is my best work, and I wrote it very quickly and distracted-ly, and I’m still mildly concerned that I’m misunderstanding the current state of string theory academia, but none of the comments have convinced me of that so far.
If you want to judge people for upvoting dubiously-accurate off-the-cuff rant posts, then … that’s fair. I complain about that too. Especially the fact that I often get far more readers on those kinds of posts than the meticulously-researched posts that I spend 4 weeks writing. Oh well, I’m not sure there’s much to be done about it. I definitely endorse not putting too much stock in lesswrong karma scores.
the author isn't really acknowledging how strong of a prior scientific consensus is
I think scientific consensuses are usually right, and sometimes overlooking big important things, and sometimes outright wrong, and we can dive into the relevant factors upstream of that. But maybe it’s not worth discussing in the abstract; it’s probably more productive to just argue about things like “should we believe X?”, and then “what scientists think about X” can be one important part of that conversation.
My worry (gut-worry?) is that cabal vs externally-opaque meritocracy is a false dilemma and cabal (for the given examples) is a strawman (ergo, opaque meritocracy).
I'm (for now, as I explicitly stated) not judging people for upvoting, I'm judging the votes! I generally can't interpret them pretty well (especially on a small scale, where I can't tell how many put how much), but I'll certainly take note if you say they're unreliable.
Thank you for the reasonable take on scientific consensus.
Fast-forward to today. There are still lots of string theorists, but now they have to work alongside a decent number of string-theory-naysayers working on directions that definitely won’t work, for reasons that members of the externally-opaque meritocracy of actually-good theoretical fundamental physicists understand well.
As an ex-String-Theorist, I'm actually pretty happy that work is also being done on, say, Loop Quantum Gravity. Now, I still have a sneaking suspicion that the Loop Quantum Gravity folks are eventually going to discover that their theory was always String Field Theory wearing a disguise (or to use the terminology, that a sector of it was dual to a sector of String Theory), even though that would require them to discover it has a critical dimension of 26 and conformal symmetry, which so far hasn't happened. But then, their theory also remains very hard, and work is ongoing (e.g. on trying to locate the ground state). As and when the LQG folks can actually construct a ground state, I think it might be rather illuminating to do String Theory as a perturbative expansion around that background spacetime.
Similarly, for another of the "competitors", the Asymptotic Freedom approach to quantum gravity, even if it isn't actually true, there is clearly a scale a little above the Planck scale where it will be a valid effective theory.
So quite a few of these "competitors" may turn out to be useful allies, or other useful ways of looking at the truth, even if String Theory turns out to be true – I don't regard most of them as obviously doomed or a waste of grant money. And it's interesting how much and where the different attempts to build a quantum theory of gravity rhyme.
Wasn't original MIRI supposed to be the "Alignment Researchers Who Are Actually Really Good At It"? Or at least that was my impression of what they were trying to do. It seems the strategy has changed.
Trying again seems worthwhile, but I think the whole point of a group like this is that enough people have a reason to hold their opinions in higher regard than others. Other fields like string theory have the older academic institutions to piggy-back off of. Alignment theorists (at least MIRI affiliated ones) seem to have made the mistake (in my opinion) of not going through academia, and I think that's been a detriment in how seriously they're taken by the public. I think this not an easy problem to fix given how much time we have left, and I also don't have any great ideas here.
Cool post!
There are still lots of string theorists, but now they have to work alongside a decent number of string-theory-naysayers working on directions that definitely won’t work, for reasons that members of the externally-opaque meritocracy of actually-good theoretical fundamental physicists understand well.
Could you elaborate on why the non-string-theory approaches can't work? You only explicitly argued for about why string theory seems like a promising approach, which is a weaker claim.
As far as I understand, the other approaches fail for reasons which are hard to describe without going into string-theoretical details. However, I have an issue with the APA coup portrayal which IMO is a part of the broader process during which postmodernism rose to popularity across the Western world and bore fruits like the ones described in Cynical Theories.
I think it’s just incorrect. Loop quantum gravity seems to be more promising and making better progress. And it is trying to directly address issues related to quantization of space-time, whereas the classical forms of string theory did their best to side-step those issues.
But I don’t know if the author of the post wants an extensive object-level debate on this here.
My impression of high-energy theorists today is that while a small number of people could be clearly identified as string-theory-naysayers, many more researchers cannot be cleanly categorized into string theorists or naysayers. Lots of people have a background in string theory, and do string-theory-inspired work trying to understand theoretical issues in quantum field theory, without believing that string theory describes the real world. They may identify as "string theorists" for social reasons and to signal what set of shared knowledge their research program assumes, but they're not directly working on the goal of "find a correct theory of quantum gravity".
An example of what I'm trying to get at: https://4gravitons.com/2014/08/22/am-i-a-string-theorist/
As to why people are interested in general questions in quantum field theory beyond the particular quantum field theories that actually describe reality, I think of it as analogous to Lagrange's and Hamilton's research into analytical mechanics frameworks beyond the particular forces that govern Newtonian physics. (I think) the hope is that understanding better what quantum field theories are, and what they can be, would eventually aid in the construction of new theories that do describe the real world.
How similar is this dynamics to Yudkowsky's point: "Alignment mindset is fundamentally difficult to obtain for a project because Graham's Design Paradox applies. People with only ordinary paranoia may not be able to distinguish the next step up in depth of cognition, and happy innocents cannot distinguish useful paranoia from suits making empty statements about risk and safety. They also tend not to realize what they're missing. "
Thanks for this post. I think the problem it names is a real bottleneck for field building and grant making in AI safety rather than only an epistemics puzzle (which, as I read them, was the focus of many comments).
First, the question: "Is this group an opaque meritocracy or a cabal?" often isn't well-posed globally. 'Consensus' underlying meritocratic transparency is tied to a specific scientific culture and depends on its granularity. For example, the goals, standards of evidence, and scientific values of 'physics' as a whole are not the same as for 'high-energy physics' as a whole or 'quantum gravity'. I like string theory as a reference for AI safety because it did involve the explicit collaboration between physicists and mathematicians that led to the formulation of a combined scientific culture (in string theory). Clearly, this led to concerns over the health and direction of physics as a field, which is echoed in the many opinions for what counts as pursuit-worthy new directions in AI safety. Also relevant to AI safety is that the scientific cultures that populate a field tend to get more granular over time as we learn more, and that we have a unique (if less 'natural', in an academic sense) field building opportunity to decide how to incorporate new ideas. I hope we can do this in a way that does not result in a collection of mutually opaque meritocracies. Personally, I think an approach that is pluralist about research programs and strict about the artifacts that should come out of those programs strikes a good balance, though defining this merit criteria per subculture in a way that aggregates to a comprehensive 'good' collective is hard.
This leads to a second point: this criteria changes over time. In a rough historical sketch, the String Wars came after a first revolution (mid 1980s work fixing anomalies, resulting in a landscape of superstring theories as QG candidates), a period of criticism (for example, in 1986: desperately seeking superstrings) and revolution again (1990s work on dualities and introducing branes, AdS/CFT). As Dmitry pointed out, this had implications on the field's scientific culture, including its goals and normative standards. I haven't thought this analogy through, but this may have a parallel in AI safety, with 'string theory as quantum gravity' standing in epistemically for agent foundations, and the empirical contact of AdS/CFT with real-world complex systems playing the role of mechanistic interpretability.
Lastly, the string wars example is as much about resource scarcity/allocation as criteria for 'good science'. A less granular example is the Anderson-Weinberg fight over the SSC, where each subfield was unable to adjudicate the other's merit while competing for the same funding. It's also worth noting that resource competition is not only where consensus gets expressed, but also how it solidifies within a scientific culture. This is perhaps another lesson for AI safety. Our scarce resource may not be less about money than grantmaker attention and limits to individual expertise (so that 'opaque' defaults to 'unfundable'). I agree with Jonas' comment below.
When I read Brian Greene's "The Elegant Universe", I kept thinking "huh, this sure reminds me of alignment!" Both in terms of social dynamics, motivations, poor feedback loops, focus on mathematical elegance, and so on. Consuming the book made me much more sympathetic to string theory than before, when my (stringy) diet consisted mainly of Woit's blog. This, I think, is common in science. The preponderance of rivalrous paradigms are always much closer than you think--if they weren't, there'd be no rivalry! Like the calorie theory of heat vs the kinetic theory of heat, which both gave quantitative predictions that matched the data of their most rigorous experiments at the time.
Anyway, all of this is to say that I wound thinking that alignment researchers were like string theorists, and vice versa. For which group is this an unflattering comparison? Well, I leave that as an exercise to the reader.
There's one key difference, and it's that for the string theory case, we want to go beyond the effective field theories of general relativity and QFT/Standard Model of particle physics, and in particular they're not aiming to control a system, but to understand how the universe works.
(One could validly argue that even accepting that string theory is promising/likely for our own universe, that the strategy to testing string theory on whether it describes the universe accurately should be radically different than current string theory research.)
But for AI alignment, we are aiming to control/engineer something, with science necessary as a byproduct of unique features of AI risk, and it's plausible we might only need an effective field theory equivalent (but stated in much vaguer terms than physics does), and humans might not need to understand the limiting cases/the cases where the effective field theory fails, because we can outsource those problems to AIs.
(And this is of course very, very deeply debated, with fractal cruxes like timelines, takeoff speeds, how good AI/humans needs to be at domains where things are hard to verify, how well AI control/AI automated alignment works and more, and notably some of the cruxes can be traded off because you don't need to get everything right, and wrongness can be tolerated, so long as it's limited (but how limited we are is yet another deep crux here), so I won't rehash these debates here)
Put another way, even accepting that the Agent Foundation worldview is correct at the limits of power and efficiency of real-world AIs that we can build in the long-term, it's possible to build AIs that are far enough from those limits of power and efficiency such that the effective theories/approaches that AI safety people like Redwood Research can work.
This is unlike string theory or other quantum gravity approaches, where we know that the effective theories we have do not work to lead us anywhere close to testing them (with one caveat.)
That's fair, but when I said "alignment" I was thinking about agent foundations in particular. Which does focus more on understanding most prosaic alignment work.
Yeah, this really does feel like the modern version of the mostly unproductive AGI/ASI/superintelligence debates, where people have implicitly different bars for what counts and doesn't count for these concepts, and this was due to different empirical beliefs that were only somewhat validated for people thinking about AI seriously, because things that were assumed to be bundled came apart.
This seems to suggest that when voting is involved, the median voter is extremely important. Cynically, we should expect things that the median voter does not understand to either go extremely wrong, or to go right for reasons that have nothing to do with understandings, but rather things like "the median voter still blindly believes the people with the right credentials, who happen to be sincere on this topic".
Also makes me wonder who is the median member of the rationalist community.
A lot of this happened before I was old enough to pay attention, but IDK if the string theorists ever presented a unified front to say, hey, we know what we're doing is more math than physics. We have strong reasons to believe that we, as a civilization, as not ready to take the next big step in theoretical physics, due to inadequate mathematical tools. We're nevertheless too applied for the pure math departments, and too narrowly focused and pure for the applied math departments. Should we split the theoretical physics departments into fundamental theoretical physics and [some other term] until that changes?"
Likewise, I suspect that the analytic philosophers of 1979 could trounce the continental philosophers at, I dunno, externally-legible tests vaguely related to philosophical reasoning, like the LSAT. If so, the rebels didn’t care.
Maybe an example is the use of logic in computer science.
I keep thinking the deep theory of modal logic developed in philosophy departments should have some impact on math, but it doesn't seem to have happened yet, as far as I know.
Modal logics are just fragments of classical first order logic (there are clever ways of doing this, illustrated in modal logics textbooks but there is also a dumb systematic way of doing it, if you're comfortable with many-sorted logic).
As such, I don't expect it to be particularly innovative in deep novel ways mathematical logicians haven't considered.
But I know of some uses of modal logic in constructive math which simplifies (the presentation of) some forcing arguments for example.
I also personally use modal logic to more easily remember Tarski-Vaught's test, which actually is the converse Barcan formula in disguise.
They're decidable fragments.
It seems to me that should make a big difference, but apparently that doesn't count for as much as I thought, because symbolic computation (also really a collection of decision procedures) hasn't lived up to the grand promise that Zeilberger sold us. (A modest example: computer algebra systems have been used to do Apéry's proof of the irrationality of ζ(3), but despite Zeilberger's efforts I can't think of any big irrationality proofs that were discovered with symbolic computation.)
Classical propositional logic is also decidable, but it's not very useful in mainstream math because it's not very expressive (which doesn't mean not useful at all. I'm aware of at least one application in graph theory).
And with all my respect to Zeilberger, he doesn't do very "classical" mathematical logic. Which is not bad at all! But you can't rely on someone going against the grain to make an opinion on the general direction of a field (which, on an unrelated note, is what I suspect this essay basically is).
If you want stuff about irrationality, there's deep work connecting Schanuel's conjecture to logic. I don't know if that body of knowledge can currently produce concrete irrationality proofs, but really only because of my lack of expertise.
Maybe I didn't get across the connection to Zeilberger, it's not that he's doing or opining about logic, it's that he's developing decision procedures, so the disappointing yield from that research program suggests that maybe decision procedures aren't so important. Like, reaching to something more basic than Zeilberger's work for an example, there's Gosper's algorithm—Gosper's paper was called "Decision procedure for indefinite hypergeometric summation". Symbolic computation in the Zeilberger style is developing decision procedures for narrow classes of mathematical questions. The connection to irrationality isn't about logic, it's Zeilberger's thought that since we have decision procedures for questions about hypergeometric sums, we should be able to use them somehow to search for hypergeometric sums to use in irrationality proofs (inspired by how Apéry used handmade hypergeometric sums to prove irrationality of ζ(3)).
Sure, but pretty much all of applied model theory is about developing 'decision procedures' for certain classes of structures. That's as central to mathematical logic as prime numbers are to arithmetic.
Establishing the decidability of a first order theory and/or quantifier-elimination are the bread an butter of the discipline. Sometimes it works, sometimes it doesn't and sometimes it's still an open problem.
Maybe look up Tarski's proof of Hilberth's 17th problem? Or the model-theoretic proof of the Nullstellensatz? These are the prototypical examples of how it goes when it goes well.
You know, it occurs to me that the Barcan formula is by Ruth Barcan Marcus, who is mentioned in the post, in the Dennet quote:
There were nominating speeches and rebuttals, the most memorable of which was by Ruth Marcus, whose Yale colleague John Smith, a philosopher of religion and a theologian, was the pluralists’ candidate. She explicitly trashed his whole career, his character, his books. I had never heard a philosopher speak so ill of a colleague in public, and seldom in private.
So not just a (minor) example of the significance of modal logic from the perspective of mathematics, but from specifically the person Dennet was talking about.
Entirely coincidental. I haven't gotten to the analytic philosophy part yet, even!
To make that example slightly less minor, because (in the appropriate modal logic etc.) the hypotheses of the Tarski-Vaught test as well as its conclusion are statements in that logic (and thus in first order logic), that result can be given a purely syntactic proof, which happens to be much nicer to visualize than the usual ugly inductive argument (in short, you just move the square/diamond for left to right in the formula).
The Barcan formula (in that specific modal logic) also happens to be a tautology, for what it's worth.
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that you disagree with is:
You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill.
…Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue.
(“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”)
…And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what’s really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out of town.
Here are two historical anecdotes, which I personally think exemplify this dynamic. But other people will say “No, one or both of those is actually the first bullet point!!” And I probably won’t be able to convince them that they’re wrong. Ironic! But that’s exactly the point!
(1) The breaching of the string theory consensus in the 2000s.[1]
“The Standard Model of Particle Physics including weak-field quantum general relativity (GR)” (I wish it were better-known and had a catchier name) appears sufficient to explain everything that happens in the solar system (ref). Nobody has ever found any experiment violating it, despite extraordinarily precise and varied tests. This theory can’t explain everything that happens in the universe—in particular, it can’t make any predictions about either (i) microscopic exploding black holes or (ii) the Big Bang. Also, (iii) the Standard Model happens to include 18 elementary particles (depending on how you count), because those are the ones we’ve discovered; but the theoretical framework is fully compatible with other particles existing too, and indeed there are strong theoretical and astronomical reasons to think they do exist. It’s just that those other particles are irrelevant for anything happening on Earth—so irrelevant that we’ve spent decades and billions of dollars searching for any Earthly experiment whatsoever where they play a measurable role, without success. Anyway, I think there are strong reasons to believe that our universe follows some set of orderly laws—some well-defined mathematical framework that elegantly unifies the Standard Model with all of GR, not just weak-field GR—even if physicists don’t know what those laws are yet. So it would be cool to find those orderly laws. This is the goal of fundamental physics theory research.[2]
Now, the extraordinary success of “The Standard Model of Particle Physics including weak-field quantum GR” is a double-edged sword. It’s good because it seems like we can already correctly predict the results of every fundamental-physics lab experiment that we’ve ever done or know how to do, which is very cool! It’s bad because we know that we don’t yet have the last word on how the universe works, but if we want to go further, and find those orderly laws mentioned above, we probably won’t be able to show off any experimental proof that we’re on the right track. Every experimental result is already covered! There’s nothing left!
So that’s the background context. Now in the 1980s–2000s, theoretical fundamental physicists investigated string theory with growing interest. They found that string theory had the intriguing property that, if a universe was governed by this kind of theory, (1) it would exactly satisfy GR (not just weak-field GR), always and everywhere, and (2) in the limit where our universe has the Standard Model of Particle Physics, a string theory universe would … well, it wouldn’t necessarily be exactly the Standard Model of Particle Physics, but at least it would be that kind of a thing, namely a “Quantum Field Theory”. And, (3) it would be well-defined everywhere, offering answers (in principle) even to questions about microscopic exploding black holes and so on, questions on which existing physics theory would return “[error: invalid input]”. None of that proves that our universe is governed by something like string theory, of course. But it sure suggests that it’s a promising research direction. So by the 2000s, the theoretical fundamental physics departments were full of string theory experts, judging the promise of each other’s string-theory papers along metrics that outsiders could not understand. Alas, this area of math / physics remains very hard, with many open questions, so as of today we still have neither a specific string theory model that exactly matches our universe, nor any reason to believe that one doesn’t exist. Work is ongoing.
Anyway, the string theorists, lacking any proof of progress that any moron could be impressed by, had formed into (what I think is) an externally-opaque meritocracy. However, others thought that it was a self-dealing cabal enmeshed in groupthink. So they led an anti-string-theory rebellion, which got a ton of press, especially through widely-discussed pop-science books by by Peter Woit and by Lee Smolin in 2006, saying that string theory was a failure and waste of money.[3]
My vague impression is that this rebellion won a lot of converts in the general public, and a smaller but still significant number of converts even within physics departments. The converts were not the existing theoretical fundamental physicists, of course, but it turns out that physics departments are pretty broad, and so they also contain experimental physicists, and experts on thermodynamics of liquid crystals, and many other niches of people who don’t necessarily understand string theory, or why people were working on it, or why the string theorists were [rightly] skeptical of alternative approaches in the literature. And some of these insider converts to the rebellion were physics department chairs, or on hiring committees, or in charge of government physics grant allocation decisions, etc.
Thus, the string theory blockade was broken.
Fast-forward to today. There are still lots of string theorists, but now they have to work alongside a decent number of string-theory-naysayers working on directions that definitely won’t work, for reasons that members of the externally-opaque meritocracy of actually-good theoretical fundamental physicists understand well.
(2) The breaching of an analytic-philosophy consensus in 1979
This amusing anecdote comes from Dan Dennett’s memoir:
Afterword
A related mental model[4]
You can (A) judge ideas via new external evidence, and/or (B) judge ideas via internal discernment of plausibility, elegance, self-consistency, consistency with already-existing knowledge and observations, etc. There’s a big range in people’s ability to apply (B) to figure things out. But what happens in “normal” sciences like biology is that there are people with a lot of (B), and they can figure out what’s going on, on the basis of hints and indirect evidence. Others don’t. The former group can gather ever-more-direct and ever-more-unassailable (A)-type evidence over time, and use that evidence as a cudgel with which to beat the latter group over the head until they finally get it. (“If you don’t believe my 7 independent lines of evidence for plate tectonics, OK fine I’ll go to the mid-Atlantic ridge and gather even more lines of evidence…”)
This is an important social tool, and explains why bad scientific ideas can die, while bad philosophy ideas live forever. And it’s even worse than that—if the bad philosophy ideas don’t die, then there’s no common knowledge that the bad philosophers are bad, and then they can rise in the ranks and hire other bad philosophers etc. Basically, to a first approximation, I think that mainstream human institutions are not really up to the task of making intellectual progress systematically over time, except where idiot-proof verification exists for that intellectual progress (for an appropriate definition of “idiot”, and with some other caveats).
…And another mental model
In The Median Researcher Problem, John Wentworth suggests that fields can sustain a healthy meritocracy approximately when a median researcher can judge good work from bad, and [John forgot to add] actually cares and is motivated and empowered to publicly praise the good stuff and rip apart the bad stuff. Here I’m suggesting that string theory and analytic philosophy were doing OK on those metrics. But then the rebels effectively expanded the pool of people whose opinions mattered, thus dragging down the median and collapsing the equilibrium.
Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains?
Alas, the answer seems to be basically “no”.
The string theorists kept proposing lots of fascinating, deep, non-obvious math conjectures, and then the mathematicians kept proving that those conjectures were true. Did that give pause to Woit, Smolin, and the other anti-string-theory agitators? Nope! They brushed it aside with a one-sentence retort: “Well so what, if you’re doing good math, then go get math grants, I’m arguing that you’re a failure at physics!”
Likewise, I suspect that the analytic philosophers of 1979 could trounce the continental philosophers at, I dunno, externally-legible tests vaguely related to philosophical reasoning, like the LSAT. If so, the rebels didn’t care.
And I mean, it’s not even that irrational to not care! It’s totally possible for people to be great at one domain but completely wrong about some domain that seems adjacent. For example, Geoff Hinton and Yann LeCun both have stellar and well-deserved reputations as AI experts, but they’re diametrically opposed on the issue of AI extinction risk, which means that at least one of them is catastrophically wrong. Speaking of which…
This post is secretly about superintelligent AI, isn’t it?
Yup! If you can reason well about what’s gonna happen with superintelligence, then that just doesn’t help you create easily-verifiable artifacts that any moron can recognize to be proof that you are correct on the specific questions at issue. Maybe you can try to get credibility in adjacent, more-verifiable domains, but per above, we should expect that to be only slightly helpful.
So, let’s suppose we follow @johnswentworth’s advice to form a Superintelligence Alignment Researchers Who Are Actually Really Good At It Club, and we reject 99% of applicants to our club,[5] based on existing club members aggressively scrutinizing the skills and judgment of each new applicant.
If we do that, at best we can create an externally-opaque meritocracy. And we will be vulnerable to attack from outsiders, who will say that our club is not in fact an externally-opaque meritocracy, but rather a self-dealing cabal enmeshed in groupthink. And those outsiders, right or wrong, will be widely believed.
(Maybe our meritocracy-or-perhaps-cabal can get funding, from the scattered people who believe us. And maybe we can make great intellectual progress within our little bubble. But getting broad acceptance of our ideas, in a way that’s obvious to outsiders, is a different matter entirely.)
Of course, I’m not saying anything here that isn’t already very obvious. Nor do I have any great solution ¯\ˍ(ツ)ˍ/¯ The not-great solution is to cross our fingers and hope that a sufficient number of key decisionmakers have the time and talent to wade into the object-level debates and correctly figure out who is right on questions of superintelligence alignment. (Or hope that the key decisionmakers follow their local vibes and memes, and wind up trusting the right people by coincidence.) Anyway, our meritocracy-or-perhaps-cabal can of course help on the margin by doing good pedagogy, outreach, etc.
I’m confident in what I say about physics, but less confident in the history-of-physics parts. I wasn’t directly involved (I was a student in physics academia at the time, but not fundamental physics), and am going off vague memories of stuff I read or heard. People can correct me. In other news, the first paragraph of this section is partly copied from here.
No comment on whether pursuing this goal is a good use of expert time and taxpayer money. That wasn’t the issue; people on both sides of the string wars were in favor of pursuing this goal.
An additional factor, which built on the momentum from Woit, Smolin, and others, was the failure in the early 2010s of the Large Hadron Collider (LHC) to find “superpartner” particles. String theory makes a strong prediction that superpartners exist. (Indeed, even without string theory, there would be strong reasons to believe that superpartners exist!) String theory did not make a strong prediction that superpartners would definitely be found at the LHC; rather, it predicted that superpartners might be found at the LHC. The best string theorists were well aware of this fact, but the meme that the LHC would imminently prove or disprove string theory nevertheless spread widely, for lots of reasons including: (1) anti-string-theory people spread this idea out of ignorance and/or malice, and/or (2) some string theorists were myopically over-hyping their latest speculative theoretical models, and/or (3) everyone wanted to convince politicians to keep funding the LHC, and/or (4) nuance was lost in a game of telephone. Anyway, the LHC turned on, no superpartners were found, and many people were upset that the string theorists were not admitting defeat.
(There’s an obvious parallel with how people in my own field of Artifical General Intelligence (AGI) safety have been trumpeting the possibility that AGI might happen very very soon. To be clear, maybe trumpeting that possibility is the right thing to do, all things considered, but we should obviously be braced for a backlash.)
This section is mostly copied from §1.4 here.
I say “our club” because I strongly believe John would let me in. Or if not, screw him, I’ll start my own club.