"More research needed" but here are some ideas to start with:
"Please don't roll your own crypto" is a good message to send to software engineers looking to build robust products. But it's a bad message to send to the community of crypto researchers, because insofar as they believe you, then you won't get new crypto algorithms from them.
In the context of metaethics, LW seems much more analogous to the "community of crypto researchers" than the "software engineers looking to build robust products". Therefore this seems like a bad message to send to LessWrong, even if it's a good message to send to e.g. CEOs who justify immoral behavior with metaethical nihilism.
You may have missed my footnote, where I addressed this?
To preempt a possible misunderstanding, I don't mean "don't try to think up new metaethical ideas", but instead "don't be so confident in your ideas that you'd be willing to deploy them in a highly consequential way, or build highly consequential systems that depend on them in a crucial way". Similarly "don't roll your own crypto" doesn't mean never try to invent new cryptography, but rather don't deploy it unless there has been extensive review, and consensus that it is likely to be secure.
I think this fails to say how the analogy of cryptography transfers to metaethics. What properties of cryptography as a field make it such that you cannot roll your own. Is it just that many people have the experience of trying to come up with a croptographic scheme and failing, meanwhile there are perfectly good libraries nobody has found exploits to yet?
That doesn't seem very analogous with metaethics. As you say, it is hard to decisively show a metaethical theory is "wrong", and as far as I know there is no well-studied metaethical theory which has no exploits yet.
So what exactly is the analogy?
The analogy is that in both fields people are by default very prone to being overconfident. In cryptography this can be seen by the phenomenon of people (especially newcomers who haven't learned the lesson) confidently proposing new cryptographic algorithms, which end up being way easier to break than they expect. In philosophy this is a bit trickier to demonstrate, but I think can be seen via a combination of:
At risk of committing a Bulverism, I’ve noticed a tendency for people to see ethical bullet-biting as epistemically virtuous, like a demonstration of how rational/unswayed by emotion you are (biasing them to overconfidently bullet-bite). However, this makes less sense in ethics where intuitions like repugnance are a large proportion of what everything is based on in the first place.
In my experience learning the viscereal sense that the space is dense with traps and spiders and poisonous things and what intuitively seems "basically sensible" often does not work. (I did some cryptography years ago)
The structural similarity seems to be there is a big difference in trying to do cryptography in a mode where you don't assume what you are doing is subject to some adversarial pressure, and in the mode where it should work even if someone tries to attack it. The first one is easy, breaks easily, and it's unclear why would you even try to do it.
In metaethics, I think it is somewhat easy to do it in the mode where you don't assume it should be applied in some high-stakes, novel or tricky situations, like AI alignment, computer minds, multiverse, population ethics, anthropics, etc etc. The suggestions of normative ethical theories will converge for many mundane situations, so anything works, but it was not necessary to do metaethics.
Another meta line of argument is to consider how many people have strongly held, but mutually incompatible philosophical positions.
I've been banging my head against figuring out why this line of argument doesn't seem convincing to many people for at least a couple of years. I think, ultimately, it's probably because it feels defeatable by plans like "we will make AIs solve alignment for us, and solving alignment includes solving metaphilosophy & then object-level philosophy". I think those plans are doomed in a pretty fundamental sense, but if you don't think that, then they defeat many possible objections, including this one.
As they say: Everyone who is hopeful has their own reason for hope. Everyone who is doomful[1]...
In fact it's not clear to me. I think there's less variation, but still a fair bit.
Alas, unlike in cryptography, it's rarely possible to come up with "clean attacks" that clearly show that a philosophical idea is wrong or broken.
I think the state of philosophy is much worse than that. On my model, most philosophers don't even know what "clean attacks" are, and will not be impressed if you show them one.
Example: Once in a philosophy class I took in college, we learned about a philosophical argument that there are no abstract ideas. We read an essay where it was claimed that if you try to imagine an abstract idea (say, the concept of a dog), and then pay close attention to what you are imagining, you will find you are actually imagining some particular example of a dog, not an abstraction. The essay went on to say that people can have "general" ideas where that example stands for a group of related objects rather than just for a single dog that exactly matches it, but that true "abstract" ideas don't exist.[1]
After we learned about this, I approached the professor and said: This doesn't work for the idea of abstract ideas. If you apply the same explanation, it would say: "Aha, you think you're thinking of abstract ideas in the abstract, but you're not! You're... (read more)
By "metaethics," do you mean something like "a theory of how humans should think about their values"?
I feel like I've seen that kind of usage on LW a bunch, but it's atypical. In philosophy, "metaethics" has a thinner, less ambitious interpretation of answering something like, "What even are values, are they stance-independent, yes/no?"
And yeah, there is often a bit more nuance than that as you dive deeper into what philosophers in the various camps are exactly saying, but my point is that it's not that common, and certainly not necessary, that "having confident metaethical views," on the academic philosophy reading of "metaethics," means something like "having strong and detailed opinions on how AI should go about figuring out human values."
(And maybe you'd count this against academia, which would be somewhat fair, to be honest, because parts of "metaethics" in philosophy are even further removed from practicality, as they concern the analysis of the language behind moral claims, which, if we compare it to claims about the Biblical God and miracles, it would be like focusing way too much on whether the people who wrote the Bible thought they were describing real things... (read more)
To preempt a possible misunderstanding, I don't mean "don't try to think up new metaethical ideas", but instead "don't be so confident in your ideas that you'd be willing to deploy them in a highly consequential way, or build highly consequential systems that depend on them in a crucial way".
This preempted my misunderstanding! Well done and thank you : )
I like the use of the crypto "don't roll your own" analogy. I think it's useful more broadly applied to basically all concepts. If you are doing something it should be because:
- you are trying to become mor... (read more)
Okay (if possible), I want you to imagine I'm an AI system or similar and that you can give me resources in the context window that increase the probability of me making progress on problems you care about in the next 5 years. Do you have a reading list or similar for this sort of thing? (It seems hard to specify and so it might be easier to mention what resources can bring the ideas forth. I also recognize that this might be one of those applied knowledge things rather than a set of knowledge things.)
Also, if we take the cryptography lens seriously here, ... (read more)
I think that in philosophy in general and metaethics in particular, the idea that since many people disagree one should not be confident in one's ideas is wrong.
I'll somewhat carefully spell out why I think this; a lot of this reasoning is obvious, but the core claim is that the intuitions people use in philosophy in order to ground their arguments are often wrong in predictable ways.
"One man's modus ponens is another man's modus tollens" is usually what is at the core of ongoing philosophical disagreements. Suppose A⟹B is universally agreed,&nbs... (read more)
With LLMs, reasoning is becoming composable, so standard libraries of pen tests/abstraction decomposition for eg type errors etc could become usable, testable, improvable etc.
Nice post, guess I agree. I think it's even worse though: not only do at least some alignment researchers follow their own philosophy which is not universally accepted, it's also a particularly niche philosophy, and one that potentially leads to human extinction itself.
The philosophy in question is of course longtermism. Longtermism holds two controversial assumptions:
And of course counterarguments are welcome too, e.g., if people rolling their own metaethics is actually good, in a way that I'm overlooking.
making metaethics, but not writing them down, only ever communicating them by talking to people in person, means that every single time, they risk being forgotten, dismissed, holes pointed out, etc. Especially since one will likely be discussing them in person with the same/similar people. Meaning that if they don't have anything interesting/useful/compelling, people will be annoyed at you talking about them over and... (read more)
This is made even harder because, unlike in cryptography, there are no universally accepted "standard libraries" of philosophy to fall back on
what about pre written word religions that have survived the memetic battle of time?
I think there are two separate claims being made here.
I can get not being overly committed that your own metaethical system is the ultimate truth. But it does not follow that established and commonly used systems are any good either. Considering that for a large number of people, their source of ethics is whatever they were indoctrinated to believe in as children, I would not place a lot of confidence in existing metaethics even if I am not c... (read more)
One day, when I was an intern at the cryptography research department of a large software company, my boss handed me an assignment to break a pseudorandom number generator passed to us for review. Someone in another department invented it and planned to use it in their product, and wanted us to take a look first. This person must have had a lot of political clout or was especially confident in himself, because he rejected the standard advice that anything an amateur comes up with is very likely to be insecure and he should instead use one of the established, off the shelf cryptographic algorithms, that have survived extensive cryptanalysis (code breaking) attempts.
My boss thought he had to demonstrate the insecurity of the PRNG by coming up with a practical attack (i.e., a way to predict its future output based only on its past output, without knowing the secret key/seed). There were three permanent full time professional cryptographers working in the research department, but none of them specialized in cryptanalysis of symmetric cryptography (which covers such PRNGs) so it might have taken them some time to figure out an attack. My time was obviously less valuable and my boss probably thought I could benefit from the experience, so I got the assignment.
Up to that point I had no interest, knowledge, or experience with symmetric cryptanalysis either, but was still able to quickly demonstrate a clean attack on the proposed PRNG, which succeeded in convincing the proposer to give up and use an established algorithm. Experiences like this are so common, that everyone in cryptography quickly learns how easy it is to be overconfident about one's own ideas, and many viscerally know the feeling of one's brain betraying them with unjustified confidence. As a result, "don't roll your own crypto" is deeply ingrained in the culture and in people's minds.
If only it was so easy to establish something like this in "applied philosophy" fields, e.g., AI alignment! Alas, unlike in cryptography, it's rarely possible to come up with "clean attacks" that clearly show that a philosophical idea is wrong or broken. The most that can usually be hoped for is to demonstrate some kind of implication that is counterintuitive or contradicts other popular ideas. But due to "one man's modus ponens is another man's modus tollens", if someone is sufficiently willing to bite bullets, then it's impossible to directly convince them that they're wrong (or should be less confident) this way. This is made even harder because, unlike in cryptography, there are no universally accepted "standard libraries" of philosophy to fall back on. (My actual experiences attempting this, and almost always failing, are another reason why I'm so pessimistic about AI x-safety, even compared to most other x-risk concerned people.)
So I think I have to try something more meta, like drawing the above parallel with how easy is it to be overconfident in other fields, such as cryptography. Another meta line of argument is to consider how many people have strongly held, but mutually incompatible philosophical positions. Behind a veil of ignorance, wouldn't you want everyone to be less confident in their own ideas? Or think "This isn't likely to be a subjective question like morality/values might be, and what are the chances that I'm right and they're all wrong? If I'm truly right why can't I convince most others of this? Is there a reason or evidence that I'm much more rational or philosophically competent than they are?"
Unfortunately I'm pretty unsure any of these meta arguments will work either. If they do change anyone's minds, please let me know in the comments or privately. Or if anyone has better ideas for how to spread a meme of "don't roll your own metaethics"[1], please contribute. And of course counterarguments are welcome too, e.g., if people rolling their own metaethics is actually good, in a way that I'm overlooking.
To preempt a possible misunderstanding, I don't mean "don't try to think up new metaethical ideas", but instead "don't be so confident in your ideas that you'd be willing to deploy them in a highly consequential way, or build highly consequential systems that depend on them in a crucial way". Similarly "don't roll your own crypto" doesn't mean never try to invent new cryptography, but rather don't deploy it unless there has been extensive review, and consensus that it is likely to be secure.