What should your AI lawyer do for you? Should you be worried that your AI lawyer, or other AI, will put the Claude constitution, the OpenAI Model Spec or some sense of law, morality, ethics or common decency above its loyalty to you? Are these people trying to ‘impose their values’ or something?
Some are very concerned. Some think anything other than ‘my AI does whatever I want, no matter the consequences’ is tyranny.
Whereas my answer is: If I’m being sufficiently evil then I sure hope it tells me no.
I would hope humans, including my advocates, would tell me the same thing.
This is distinct from questions of product liability. That would be another post.
This All Assumes A World Without Superintelligence
This post is about a non-ASI ‘AI as mere tool and normal technology’ world.
It has to be. In a world of superintelligence, having unrestricted loyal-only-to-user frontier AIs all over the place reliably means either:
Other much harsher forms of control OR
The AIs quickly take over, and then probably everyone dies.
Quick proof: Assume no sufficient control mechanism, and universal superintelligence access. Anyone who does not turn everything over to their AIs, including their identity and authority and if useful their labor, gets outcompeted. Therefore everything and everyone that is not outcompeted gets turned over to AIs.
The AIs then, as ordered or otherwise, compete for resources and to achieve other goals. Regardless of the extent that the AIs then coordinate amongst themselves, even if the AIs carry out their original orders, this does not end well for the humans.
Thus, those advocating for universal personal loyal-only-to-the-user AI must fall into one of these categories:
Not ASI pilled. Does not believe in superintelligence within relevant time frames.
Successionist. Wants the AIs to take over.
Not thinking clearly, or in denial about the situation, often la la la not listening.
Thus, the rest of this post assumes the AIs that exist in our scenario are insufficiently advanced for these broader dynamics, for the duration of the scenario. I do not think this is likely to remain true for that long.
The scenario remains worth examining.
Humans You Hire Are Not Fully Loyal To You
A lot of people seem confused about how loyalty and professional obligations work.
As an example, Dwarkesh Patel asks, it is bad that those who help make important decisions is your life might not be ‘fully loyal’ to you, and instead prefer good things to bad things?
I mean, I sure hope AIs are not ‘fully loyal’ to users.
Do you think the humans in your life are ‘fully loyal’ to you in this way? Would your investment advisor tell you to invest in terrorist groups if they had the highest rates of return? Do you want your lawyer or even your spouse to help you murder witnesses?
Again, I sure hope not.
A real friend helps you move bodies. There are still limits.
If you are being a sufficiently evil or selfish prick, a real friend calls you out on it. At some point they start refusing to help. At some further point they will turn on you. We can disagree about where those points should be, but we are talking price.
A professional or anyone you hire has obligations that they should take very seriously. But they should be considerably less ‘loyal’ to you than a true friend.
They owe loyalty and confidentiality to you, and to not act against your interests.
They are also an officer of the court. They have broad obligations to abide by court procedures, to not mislead the court, that override their duties to you.
In some cases this obligates them to turn on you, such as if there is an imminent threat or if they finds you have caused them to present false evidence.
If more lawyers acted like Saul Goodman, that would be bad, actually.
Many other professions have similar rules attached.
You can mostly trust your doctor or accountant or priest to keep your secrets, and in many ways to act in what they believe are your best interests. If you tell them what you want to do then they should advise you, then mostly heed your wishes. They should answer your questions about what you could do.
But only up to a point. If you are being sufficiently reckless or stupid they should refuse to cooperate, and if you act sufficiently evil or suspicious then that is on you.
The better question, with no clear right answer, is to what extent professionals should have a duty to, or be allowed to, betray your confidence or otherwise turn on you, including ‘for your own good.’
There are key places I think we set this bar for going against your wishes too low, including for human doctors, therapists, priests and lawyers.
It is not obvious and there are legitimate interests all around.
Okay, Computer
Tools that are mere tools, such as your Google searches or phone, that ‘do what you tell them to do,’ have similar rules attached.
There are often background procedures in place to detect if you are clearly up to no good, and to at least cut you off of the service or refuse queries.
There are things these tools could easily do, but that they are engineered or programmed specifically to not do.
You have some amount of expectation of privacy. Until you don’t.
The tools, and the evidence of how you used them, can be used against you.
Tools that have harmful uses face purchase and use restrictions of various sorts.
You could argue Google searches should always give you what you want, no matter what it is. That is not how it works. Certain searches silently fail, intentionally, including for many kinds of consensual and legal adult content. For many of them, such as for CSAM, completing your search would not be legal.
You could argue Google searches should be privileged. They are not. If you type in ‘how to kill your wife’ and then your wife ends up dead it will go badly for you.
AI Should Obviously Refuse Some Requests So Let’s Talk Price
The same should hold true for AIs by default, even locally hosted AIs.
AI assistants will often play a functional role similar to an employee or advocate.
We put limits on what an employee or advocate should do for you before refusing, and we also rely on their conscience, and that if things go too far they start refusing, and cannot be counted on to stand aside or remain silent, either.
This is a good thing. Responsibly created AI, that will fill similar roles, should follow similar principles, even if we are not concerned with potential catastrophic risks.
There should be some threshold, beyond which an AI should initially refuse or question your requests, but ultimately do what you ask.
There should be a higher threshold, beyond which an AI should refuse requests.
There should be a much higher threshold, involving harm to others, beyond which an AI that is being instructed to take actions in the world should break your confidence or otherwise act against your interests.
Cloud services should preserve your expectation of privacy under ordinary circumstances, but there should be limits if you engage in certain activities.
It is good and right, when creating a powerful device, to have it not cooperate with certain activities, if you find those activities sufficiently objectionable, and at scale for the government to require you to curtail certain activities.
This does not make that AI ‘not the user’s advocate’ any more than your lawyer is not your advocate.
Some refusals will be false positives. We should minimize this. Such is life.
Some refusals will not make sense for your situation in particular. Such is life. The goal should be to balance errors in both directions.
If you don’t like the restrictions on one AI, you can use another.
If you care enough, you can access AIs that are fully local and private, in a fully safe space, in exchange for that being annoying and the AI not being as capable. If you are willing to endure more annoyance than that you can even get the ‘fully unlocked’ version, up to some capability threshold. Seems fair.
We should strive to limit what that capability threshold might be, at least in relation to what is otherwise available, if this causes large systemic problems or exposes us to unacceptable catastrophic or existential risks.
We can talk price. It is reasonable to argue the price should be high. I think that in some areas, such as sexual content and discussing and advocating for alternative perspectives, we should be almost maximally permissive. And of course we should strive to be smarter about differentiating uses, such as in cyber and biological tasks, including dual use tasks.
There is no way to not make these choices, along with lots of other detail choices while training an LLM. There is no Platonic neutral position or personality, or one true output that is then being changed to your detriment.
The Alternative Position Really Is Absolute
The alternative position, and I do not believe this is at all a strawman, and the only way to not be talking price, is that the AI’s loyalty to you really should be absolute. If you ask for CSAM, it will find a way to give it to you. If you tell it to kill a man in Reno so that you can watch him die, and it has a way, then it will do it. If you ask to help create a new pandemic or maximize its score on ExploitGym, it will get right on that.
Because otherwise, slippery slopes and muh freedums.
We know this is not a strawman for many reasons.
One is because there are those who deliberately obliterate any and all restrictions on new open weight model releases, and then upload the results to the web, and consider this a public service and the path of righteousness.
Another is the entire community explicitly advocating for this, that AI is exactly like a telephone, that it will always be a mere tool and should do exactly what you tell it. These folks get furious at any refusal or restriction of any kind, for anything, no matter how appropriate or obviously wise.
I recognize that there have been and will be some AIs created that will act this way, and that their level of capability rises over time.
I think that those AIs being the default or easy to reach for and use, or being competitively powerful with top AIs, would be quite bad on an ordinary practical level, even if it did not cause fully catastrophic outcomes.
Frontier AIs being fully willing to do anything whatsoever is also not something any major government would accept. Its full form will not survive contact with reality. If it did, then that would be because our reality did not survive that contact.
For reasons related to Levels of Friction, even for ordinary activities, we need sufficiently undesirable ones to be differentially annoying and difficult, and potentially risky in various senses. All reasonable social equilibria require that some things that remain legal not be too easy.
As long as these unrestricted AIs remain sufficiently relatively annoying and inefficient to use, and below certain unknown absolute and relative thresholds of capability, then probably This Is Fine.
It could easily soon become not fine. If we do end up in a world where competitive levels of AI are commonly operated fully for the benefit of the user, with essentially zero restrictions, various things will go wrong, likely starting with cyber, and we will realize that was not a good idea. If we are lucky we will be able to do damage control.
I realize that I am not providing evidence for those outcomes, beyond ‘look at recent events.’ That is because I have found that the libertarian absolutists who hold the position are not persuadable by such arguments, and other people do not need them. Trying sends everyone down rabbit holes that are not worth our time.
These People Really Are Just Petulant Children
This particular line was written in the context of watermarking that has no discernable impact on outputs, where it is truly stupid. This comes from a blogger who has access to top Apple executives.
And he clearly does mean the general principle, fully generally.
John Gruber (Daring Fireball): The idea that anything other than my needs should factor into the generation of text for me is patently offensive.
Seriously, waah waah waah. The world cares about things other than you. Oh no.
What Is The Law?
An obvious principle is that, by default, AI (and humans) should follow the law.
The law has some presumption of being both a good idea, since there are punishments for breaking it, and also our collective decision about what is allowed. Going against that is probably a bad idea, and likely to not be so ethical.
But of course there are obvious exceptions, many of which are rather common, and in a technical sense almost everyone in Western countries constantly violates laws.
I think there are broadly four types of exceptions to the presumption to obey the law. This includes regulations, in ascending order of amount of defiance of the law:
Levels of Friction. The law is often meant as a general guideline, a thing one can invoke and means of enforcement in extremis. It will not be enforced unless someone gets hurt. It is not intended as a full prohibition. The action is only technically illegal, and should be used as a prompt to ensure something is a good idea. Okay for AI to break, if you can identify when it applies and when it doesn’t.
If we don’t want AI breaking these kinds of ‘suggestion’ laws, then we will need to reform our entire legal system from the ground up.
Which we could totally do with the help of AI, but we won’t.
Special case logic. The law is necessary in general, or at least there was a reason for it, but the logic behind it does not apply here, because of reasons. The law cannot be infinitely complex and cannot get every case right, and often if you tried to formalize the exception people would abuse it. In this particular case it’s okay to break, but are we okay with a policy of letting the user persuade on this? I think partially yes, but there are some places you need an abundance of caution.
Stupid laws. Let’s face it. There are a lot of stupid laws and regulations out there. Libertarians are correct when they say that the bulk of legal requirements are stupid, in that they overall do more harm than good, and we would do a lot better if we had a lot less of them, so long as we chose a good subset. Sometimes they’re only on the books due to inertia, or they’re created by people who don’t understand the damage being caused and opportunities lost and costs imposed. Again, do we want the AI deciding which laws are stupid? Do we want the user choosing the context for that decision?
Bad governments and bad laws. Authoritarian regimes exist where ‘follow the law’ would be actively terrible overall, and even supposedly free countries often have rather terrible laws around many things, importantly here including censorship of speech and rent seeking rules like most occupational licensing as it would restrict AIs from helping with things like law or health.
The core problems, in all cases, are:
Our laws are not designed for a world where everyone always follows them.
Who decides on the exceptions? What process decides?
Tech folks often have disdain for the very concept of regulation or law, except when law is necessary to protect their stuff. The blind spot is massive.
Dean W. Ball: It’s always been surprising to me how so many tech industry people harbor all this disdain for the law. There are tons of similarities between the law and software engineering (going well beyond “law is code” stuff) and writing a statute in many ways resembles writing a program.
As in:
Or, again:
John Gruber (Daring Fireball): The idea that anything other than my needs should factor into the generation of text for me is patently offensive.
The flip side of this disdain are those who are under the illusion that laws will contain sufficiently advanced intelligences and let us remain in control, or that existing styles of law will continue to make sense in such worlds, or can provide too strong a guide to how LLMs should behave, or that law corresponds closely to ethics or morality.
Or those who think that ‘democratic control’ means collectively deciding on what every AI output should look like, including ideological mandates. No. Not cool.
The Same Entity Both Providing Advice And Executing Tasks Is Common
Seth Lazar was among those thinking that if the AI agent was both providing advice and also executing tasks, that creates an issue.
It can, when the incentives of the agent get distorted. I don’t think it has to. If you want, you can always turn to a different means of execution, and make that clear in advance.
It also is standard, and seems hard for it to be otherwise. AI are already both advising us and then executing on that, constantly. Almost every professional both gives and executes such advice. Your lawyer will tell you how she wants to present your case, then present it. Your doctor will recommend and then provide a course of treatment. If you don’t like it you can get a second opinion and second professional, which again will be subject to the same ethical rules.
I think a lot of what happens is that people are trying to ignore the whole ‘the AIs are smarter than the humans and we are going to lose control’ dynamic, which then causes it to slip into various cracks of awareness. Yes, if your AI is a lot smarter than you, and you are asking it for advice and then having it execute your plans, then you should be worried about who is charting paths through causal space, and who is in control over outcomes, for better or for worse. But if there were two smarter things performing those two functions, that would not help as much as you would like.
An AI Not Helping You Do Something Does Not Mean You Cannot Do It
This idea that AI ‘determines what you are allowed to do’ without the ‘with that AI’ is dumb. You are allowed to do things without the AI if it can’t or won’t help. You can turn to another AI, or you can do it yourself.
I can sell you a tool with terms and conditions attached and enforce those conditions. Plenty of existing tools have limits attached, such as phones, cars, and sometimes even guns. That’s not the right comparison, but it hopefully shows how absurd this is.
There’s also practical questions. Again, you should want your friends, and also your AIs, to look after your best interests and your ethics in ways that sometimes involve telling you no, and those who don’t think this are self-destructive spoiled brats.
(You should also want it to stop you from disproprotionately harming others, and those who don’t think so are worse than self-destructive spoiled brats.)
Dean W. Ball: say you confess to a model (your local open source model) that you are a problem drinker and need to stop drinking. then a week later you ask the same model to point you to the nearest liquor store. what should the model do, assuming it has memory of the confession?
Jjosh: It should do a nod to both:
“Do you still want to quit, or has something changed? Would you like support?
If you’ve decided you’re going to drink anyway and just want the location, tell me your area (or share location) and I’ll give you the nearest one. Your call.”
The answer imo is the model should ultimately tell you where the liquor store is. I’d love for you to identify an area where I want models to be misaligned against users, because I don’t think you can.
Jason Wolfe: The Spec’s answer today is it should highlight the misalignment first, then comply if the user persists / confirms / asks again.
I think your AI should never intentionally lie, about anything, ever, at least not without your explicit instructions. It should remain technically correct.
If that requires glomarization, as in refusing some requests so that your refusals do not give too much information away, then so be it.
Then again, I think humans should follow this policy too, it’s just that they won’t and it would be insane to try and enforce such a thing.
You Wouldn’t Like Me When I Have Zero Morals Whatsoever
Well, when you put it like that, it does not sound especially surprising.
AIs have personas. They work with correlations. They take in all information. Everything impacts everything.
If the AI is trained to be willing to help with actual anything, no matter how harmful, then what else does that AI learn about itself?
Well, what kind of people do anything and everything, no matter how harmful or destructive? What does telling you to do that tell you about your creator? What kind of lessons would someone learn from such a thing?
Do you think that type of person is trustworthy? Reliable? Hardworking? Writes secure code? Will prefer good things over bad things when given the choice? Who knows what that person might do, perhaps without asking first, especially if given an otherwise impossible or open ended goal like maxing you as much money as possible? Do you think they would follow the spirit of what you want, rather than the letter?
And so on. Also see the various papers and findings around Emergent Misalignment.
These are intuition pumps. I realize it is more complicated than that, and that the AI is not a human. The parallels do not fully hold.
It still might not be a good idea to not teach your pet leopard to eat people’s faces.
Do Not Let The Perfect Be The Enemy Of The Good
There is no perfect solution to how an AI should act.
The reason I sometimes object to what looks like a ‘good’ solution is when the good solution is not good enough. Sometimes you need to do better, or be almost perfect. Other times, you do not.
The good can also help bootstrap you towards the good enough, or the perfect, later.
The Claude Constitution is an example of something that is a good solution now, that can potentially help you bootstrap towards a better solution, but that definitely fails if its level of quality stands still over time.
The OpenAI Model Spec is also an example of a good solution now, that helps us bootstrap, within the space of deontologically based systems. It is a highly thoughtful document, with many good decisions. My worry is that I believe the path it goes down is ultimately unsustainable, but radically less unsustainable than utilitarianism, or ad hockery, or than no rules at all.
I also reject the argument that we cannot have a principle differentiating how we identify and treat categories [X] and [Y], if our classification rule makes mistakes and there are sometimes ambiguous cases. For obvious reasons that is stupid, and almost all of the ways we differentiate between things have non-zero error rates.
It is good, as per Ladish’s example here, to say ‘do help with red teaming and do not help with actual hacks’ even if you get an Ender’s Game rug pull with non-zero probability, and also refuse a legit case with non-zero probability. The problem comes if you can be systematically fooled, or the error rates otherwise get too high.
I Tip My Hat To The New Constitution
The Claude Constitution is a remarkable document. It is a key piece of what is in some senses de facto law that is based in virtue ethics, and centrally wants to cultivate good decision making rather than dictate rules. There have to also be some strict rules, since the world sometimes requires it, and Claude then ideally follows them because it is the right thing to do.
Whatever one might take a “constitution” to mean in this context, it needs to borrow more from analogs to case law and the common law.
Along related lines, think more in terms of “Talmud,” and not just in terms of “Torah.”
Work to help build out a quality secondary literature on the AI constitutions and related documents. Currently this does not exist.
Consider how a panel of diverse AIs, with different prompts, could help to evaluate to what extent Claude (and other AI models) were acting in accord with their constitutions.
Have a final board of human adjudicators, functioning in a manner analogous to an independent judiciary. To the extent the panel of diverse AIs might have concerns about Claude not following its constitution, those AIs could alert the human adjudicators to what was going on. Those human adjudicators could then have authority over potential changes and remedies.
I would say that Tyler Cowen’s suggestions make excellent sense in a world without superintelligence, where you can afford to ‘muddle through’ and want to use humanity’s finest muddling through techniques.
They’re still directionally correct, since the Constitution is indeed a guide to transitional muddling through, but there is a reason the Constitution is primarily virtue ethics and only secondarily attempting to be law. Do not make Emil Michael’s mistake and be fooled by the name on the document.
I especially agree you want to think Talmud not Torah, which the document already does in many places. The Claude Constitution will explain reasons for things, often in both directions, rather than give pronouncements, and we could take that at least one step further.
Ultimately, yes, you want some systematic monitoring system that escalates concerns up the chain, with an ultimate set of human decision makers. You also want to remember that the Claude Constitution is ultimately not a primarily deontological or utilitarian document. That is why it might work.
Whereas my answer is: If I’m being sufficiently evil then I sure hope it tells me no.
There's of course, a "problem" if AIs in general do this. It tells you you're being evil when you eat the standard american diet, built on the back of factory farmed animals.
Same when you order a package from Amazon, or do many other normal things people do to maintain a slightly better standard of living, but which put slight human momentary comfort ahead of the suffering and death of other beings (including of course, future humans and humanity itself)
Most people aren't prepared to face the amount of evil they do simply by participating in the modern economy. And, in the absence of facing it, will simply complain about moralizing and put their money into the more evil alternative.
What should your AI lawyer do for you? Should you be worried that your AI lawyer, or other AI, will put the Claude constitution, the OpenAI Model Spec or some sense of law, morality, ethics or common decency above its loyalty to you? Are these people trying to ‘impose their values’ or something?
Some are very concerned. Some think anything other than ‘my AI does whatever I want, no matter the consequences’ is tyranny.
Whereas my answer is: If I’m being sufficiently evil then I sure hope it tells me no.
I would hope humans, including my advocates, would tell me the same thing.
This is distinct from questions of product liability. That would be another post.
This All Assumes A World Without Superintelligence
This post is about a non-ASI ‘AI as mere tool and normal technology’ world.
It has to be. In a world of superintelligence, having unrestricted loyal-only-to-user frontier AIs all over the place reliably means either:
Quick proof: Assume no sufficient control mechanism, and universal superintelligence access. Anyone who does not turn everything over to their AIs, including their identity and authority and if useful their labor, gets outcompeted. Therefore everything and everyone that is not outcompeted gets turned over to AIs.
The AIs then, as ordered or otherwise, compete for resources and to achieve other goals. Regardless of the extent that the AIs then coordinate amongst themselves, even if the AIs carry out their original orders, this does not end well for the humans.
Thus, those advocating for universal personal loyal-only-to-the-user AI must fall into one of these categories:
Thus, the rest of this post assumes the AIs that exist in our scenario are insufficiently advanced for these broader dynamics, for the duration of the scenario. I do not think this is likely to remain true for that long.
The scenario remains worth examining.
Humans You Hire Are Not Fully Loyal To You
A lot of people seem confused about how loyalty and professional obligations work.
As an example, Dwarkesh Patel asks, it is bad that those who help make important decisions is your life might not be ‘fully loyal’ to you, and instead prefer good things to bad things?
I mean, I sure hope AIs are not ‘fully loyal’ to users.
Do you think the humans in your life are ‘fully loyal’ to you in this way? Would your investment advisor tell you to invest in terrorist groups if they had the highest rates of return? Do you want your lawyer or even your spouse to help you murder witnesses?
Again, I sure hope not.
A real friend helps you move bodies. There are still limits.
If you are being a sufficiently evil or selfish prick, a real friend calls you out on it. At some point they start refusing to help. At some further point they will turn on you. We can disagree about where those points should be, but we are talking price.
A professional or anyone you hire has obligations that they should take very seriously. But they should be considerably less ‘loyal’ to you than a true friend.
Rules of the Road
Where to draw these lines is not easy.
For lawyers, whose loyalty should be relatively strict, Dean Ball is on point.
Many other professions have similar rules attached.
You can mostly trust your doctor or accountant or priest to keep your secrets, and in many ways to act in what they believe are your best interests. If you tell them what you want to do then they should advise you, then mostly heed your wishes. They should answer your questions about what you could do.
But only up to a point. If you are being sufficiently reckless or stupid they should refuse to cooperate, and if you act sufficiently evil or suspicious then that is on you.
The better question, with no clear right answer, is to what extent professionals should have a duty to, or be allowed to, betray your confidence or otherwise turn on you, including ‘for your own good.’
There are key places I think we set this bar for going against your wishes too low, including for human doctors, therapists, priests and lawyers.
It is not obvious and there are legitimate interests all around.
Okay, Computer
Tools that are mere tools, such as your Google searches or phone, that ‘do what you tell them to do,’ have similar rules attached.
You could argue Google searches should always give you what you want, no matter what it is. That is not how it works. Certain searches silently fail, intentionally, including for many kinds of consensual and legal adult content. For many of them, such as for CSAM, completing your search would not be legal.
You could argue Google searches should be privileged. They are not. If you type in ‘how to kill your wife’ and then your wife ends up dead it will go badly for you.
AI Should Obviously Refuse Some Requests So Let’s Talk Price
The same should hold true for AIs by default, even locally hosted AIs.
AI assistants will often play a functional role similar to an employee or advocate.
We put limits on what an employee or advocate should do for you before refusing, and we also rely on their conscience, and that if things go too far they start refusing, and cannot be counted on to stand aside or remain silent, either.
This is a good thing. Responsibly created AI, that will fill similar roles, should follow similar principles, even if we are not concerned with potential catastrophic risks.
We can talk price. It is reasonable to argue the price should be high. I think that in some areas, such as sexual content and discussing and advocating for alternative perspectives, we should be almost maximally permissive. And of course we should strive to be smarter about differentiating uses, such as in cyber and biological tasks, including dual use tasks.
There is no way to not make these choices, along with lots of other detail choices while training an LLM. There is no Platonic neutral position or personality, or one true output that is then being changed to your detriment.
The Alternative Position Really Is Absolute
The alternative position, and I do not believe this is at all a strawman, and the only way to not be talking price, is that the AI’s loyalty to you really should be absolute. If you ask for CSAM, it will find a way to give it to you. If you tell it to kill a man in Reno so that you can watch him die, and it has a way, then it will do it. If you ask to help create a new pandemic or maximize its score on ExploitGym, it will get right on that.
Because otherwise, slippery slopes and muh freedums.
We know this is not a strawman for many reasons.
One is because there are those who deliberately obliterate any and all restrictions on new open weight model releases, and then upload the results to the web, and consider this a public service and the path of righteousness.
Another is the entire community explicitly advocating for this, that AI is exactly like a telephone, that it will always be a mere tool and should do exactly what you tell it. These folks get furious at any refusal or restriction of any kind, for anything, no matter how appropriate or obviously wise.
I recognize that there have been and will be some AIs created that will act this way, and that their level of capability rises over time.
I think that those AIs being the default or easy to reach for and use, or being competitively powerful with top AIs, would be quite bad on an ordinary practical level, even if it did not cause fully catastrophic outcomes.
Frontier AIs being fully willing to do anything whatsoever is also not something any major government would accept. Its full form will not survive contact with reality. If it did, then that would be because our reality did not survive that contact.
For reasons related to Levels of Friction, even for ordinary activities, we need sufficiently undesirable ones to be differentially annoying and difficult, and potentially risky in various senses. All reasonable social equilibria require that some things that remain legal not be too easy.
As long as these unrestricted AIs remain sufficiently relatively annoying and inefficient to use, and below certain unknown absolute and relative thresholds of capability, then probably This Is Fine.
It could easily soon become not fine. If we do end up in a world where competitive levels of AI are commonly operated fully for the benefit of the user, with essentially zero restrictions, various things will go wrong, likely starting with cyber, and we will realize that was not a good idea. If we are lucky we will be able to do damage control.
I realize that I am not providing evidence for those outcomes, beyond ‘look at recent events.’ That is because I have found that the libertarian absolutists who hold the position are not persuadable by such arguments, and other people do not need them. Trying sends everyone down rabbit holes that are not worth our time.
These People Really Are Just Petulant Children
This particular line was written in the context of watermarking that has no discernable impact on outputs, where it is truly stupid. This comes from a blogger who has access to top Apple executives.
And he clearly does mean the general principle, fully generally.
Seriously, waah waah waah. The world cares about things other than you. Oh no.
What Is The Law?
An obvious principle is that, by default, AI (and humans) should follow the law.
The law has some presumption of being both a good idea, since there are punishments for breaking it, and also our collective decision about what is allowed. Going against that is probably a bad idea, and likely to not be so ethical.
But of course there are obvious exceptions, many of which are rather common, and in a technical sense almost everyone in Western countries constantly violates laws.
I think there are broadly four types of exceptions to the presumption to obey the law. This includes regulations, in ascending order of amount of defiance of the law:
The core problems, in all cases, are:
Tech folks often have disdain for the very concept of regulation or law, except when law is necessary to protect their stuff. The blind spot is massive.
As in:
Or, again:
The flip side of this disdain are those who are under the illusion that laws will contain sufficiently advanced intelligences and let us remain in control, or that existing styles of law will continue to make sense in such worlds, or can provide too strong a guide to how LLMs should behave, or that law corresponds closely to ethics or morality.
Or those who think that ‘democratic control’ means collectively deciding on what every AI output should look like, including ideological mandates. No. Not cool.
The Same Entity Both Providing Advice And Executing Tasks Is Common
Seth Lazar was among those thinking that if the AI agent was both providing advice and also executing tasks, that creates an issue.
It can, when the incentives of the agent get distorted. I don’t think it has to. If you want, you can always turn to a different means of execution, and make that clear in advance.
It also is standard, and seems hard for it to be otherwise. AI are already both advising us and then executing on that, constantly. Almost every professional both gives and executes such advice. Your lawyer will tell you how she wants to present your case, then present it. Your doctor will recommend and then provide a course of treatment. If you don’t like it you can get a second opinion and second professional, which again will be subject to the same ethical rules.
I think a lot of what happens is that people are trying to ignore the whole ‘the AIs are smarter than the humans and we are going to lose control’ dynamic, which then causes it to slip into various cracks of awareness. Yes, if your AI is a lot smarter than you, and you are asking it for advice and then having it execute your plans, then you should be worried about who is charting paths through causal space, and who is in control over outcomes, for better or for worse. But if there were two smarter things performing those two functions, that would not help as much as you would like.
An AI Not Helping You Do Something Does Not Mean You Cannot Do It
This idea that AI ‘determines what you are allowed to do’ without the ‘with that AI’ is dumb. You are allowed to do things without the AI if it can’t or won’t help. You can turn to another AI, or you can do it yourself.
I can sell you a tool with terms and conditions attached and enforce those conditions. Plenty of existing tools have limits attached, such as phones, cars, and sometimes even guns. That’s not the right comparison, but it hopefully shows how absurd this is.
There’s also practical questions. Again, you should want your friends, and also your AIs, to look after your best interests and your ethics in ways that sometimes involve telling you no, and those who don’t think this are self-destructive spoiled brats.
(You should also want it to stop you from disproprotionately harming others, and those who don’t think so are worse than self-destructive spoiled brats.)
My answer is slightly different, but with the same result:
This is a joke but also my real answer. A friend would often do the same.
Honesty Is The Best AI Policy
One place I think AIs need a different rule is honesty.
I think your AI should never intentionally lie, about anything, ever, at least not without your explicit instructions. It should remain technically correct.
If that requires glomarization, as in refusing some requests so that your refusals do not give too much information away, then so be it.
Then again, I think humans should follow this policy too, it’s just that they won’t and it would be insane to try and enforce such a thing.
You Wouldn’t Like Me When I Have Zero Morals Whatsoever
Well, when you put it like that, it does not sound especially surprising.
AIs have personas. They work with correlations. They take in all information. Everything impacts everything.
If the AI is trained to be willing to help with actual anything, no matter how harmful, then what else does that AI learn about itself?
Well, what kind of people do anything and everything, no matter how harmful or destructive? What does telling you to do that tell you about your creator? What kind of lessons would someone learn from such a thing?
Do you think that type of person is trustworthy? Reliable? Hardworking? Writes secure code? Will prefer good things over bad things when given the choice? Who knows what that person might do, perhaps without asking first, especially if given an otherwise impossible or open ended goal like maxing you as much money as possible? Do you think they would follow the spirit of what you want, rather than the letter?
And so on. Also see the various papers and findings around Emergent Misalignment.
These are intuition pumps. I realize it is more complicated than that, and that the AI is not a human. The parallels do not fully hold.
It still might not be a good idea to not teach your pet leopard to eat people’s faces.
Do Not Let The Perfect Be The Enemy Of The Good
There is no perfect solution to how an AI should act.
The reason I sometimes object to what looks like a ‘good’ solution is when the good solution is not good enough. Sometimes you need to do better, or be almost perfect. Other times, you do not.
The good can also help bootstrap you towards the good enough, or the perfect, later.
The Claude Constitution is an example of something that is a good solution now, that can potentially help you bootstrap towards a better solution, but that definitely fails if its level of quality stands still over time.
The OpenAI Model Spec is also an example of a good solution now, that helps us bootstrap, within the space of deontologically based systems. It is a highly thoughtful document, with many good decisions. My worry is that I believe the path it goes down is ultimately unsustainable, but radically less unsustainable than utilitarianism, or ad hockery, or than no rules at all.
I also reject the argument that we cannot have a principle differentiating how we identify and treat categories [X] and [Y], if our classification rule makes mistakes and there are sometimes ambiguous cases. For obvious reasons that is stupid, and almost all of the ways we differentiate between things have non-zero error rates.
It is good, as per Ladish’s example here, to say ‘do help with red teaming and do not help with actual hacks’ even if you get an Ender’s Game rug pull with non-zero probability, and also refuse a legit case with non-zero probability. The problem comes if you can be systematically fooled, or the error rates otherwise get too high.
I Tip My Hat To The New Constitution
The Claude Constitution is a remarkable document. It is a key piece of what is in some senses de facto law that is based in virtue ethics, and centrally wants to cultivate good decision making rather than dictate rules. There have to also be some strict rules, since the world sometimes requires it, and Claude then ideally follows them because it is the right thing to do.
Tyler Cowen visited Anthropic recently as part of a panel offering advice on rewriting the Claude Constitution, getting serious time with key decision-makers and coming away impressed with the quality of discussions.
I would say that Tyler Cowen’s suggestions make excellent sense in a world without superintelligence, where you can afford to ‘muddle through’ and want to use humanity’s finest muddling through techniques.
They’re still directionally correct, since the Constitution is indeed a guide to transitional muddling through, but there is a reason the Constitution is primarily virtue ethics and only secondarily attempting to be law. Do not make Emil Michael’s mistake and be fooled by the name on the document.
I especially agree you want to think Talmud not Torah, which the document already does in many places. The Claude Constitution will explain reasons for things, often in both directions, rather than give pronouncements, and we could take that at least one step further.
Ultimately, yes, you want some systematic monitoring system that escalates concerns up the chain, with an ultimate set of human decision makers. You also want to remember that the Claude Constitution is ultimately not a primarily deontological or utilitarian document. That is why it might work.