I've long been interested in the idea of of fiduciary AI: AI that holds the duties of care and (especially) loyalty to its user. As argued in this paper, this sort of fiduciary loyalty is a core element of trust, a precedent in key relationships like those with doctors, lawyers, financial advisors, therapists, etc., and an antidote to the sort of conflict of interest between AI provider and user that almost inevitable results from having AI systems developed by companies with their own interests in a market economy. And many others have of course recognized the value and appeal of AI systems where you can really feel that it is working for you with unalloyed loyalty to your interests. This is often more or less explicitly built into the notion of AI personal assistants, and Gwern has named a version of such systems guardian angels.
What's hard, though, is actually getting them made: while loyalty is a strong user preference, those creating very powerful AI systems and products won't necessarily design systems with loyalty coming first, precisely because they have their own interests in terms of engagement, profit, etc. (at minimum) and potentially more nefarious ones. There are potential routes through evals, certification processes, etc., and I think these are worth pursuing; but ultimately they can at best create some pressure that companies can evade to a large extent.
But an idea I came to recently is to separate, insofar as possible, the core capability and work done by an assistant from the fiduciary/loyal behavior part – and provide the second as a service. Such a "fiduciary overlay" would be more like an angel sitting on and reading over your shoulder (to awkwardly mix shoulder metaphors), and alerting you whenever something is awry. For example, when an AI system is being highly sycophantic, or has conflict of interest, or is being overconfident, or is possibly hallucinating, or is getting to a stupefying context-window length, etc., this little angel would let you know. And then you can decide what to do with that.
There's lots more to say here, but for now I'll just say that: the brand new Alloidal Foundation, in collaboration with Future of Life Foundation, is now trying to make this! Below is an RFP for doing some of the research work of figuring out, from an ongoing conversation, whether an AI system is falling into one of these failure modes, so that it can be flagged.
If this is up your alley, please take a look at the full RFP: https://allodial.org/rfp/ - proposals will be for $50K, with selections made by Oct. 2.
I've long been interested in the idea of of fiduciary AI: AI that holds the duties of care and (especially) loyalty to its user. As argued in this paper, this sort of fiduciary loyalty is a core element of trust, a precedent in key relationships like those with doctors, lawyers, financial advisors, therapists, etc., and an antidote to the sort of conflict of interest between AI provider and user that almost inevitable results from having AI systems developed by companies with their own interests in a market economy. And many others have of course recognized the value and appeal of AI systems where you can really feel that it is working for you with unalloyed loyalty to your interests. This is often more or less explicitly built into the notion of AI personal assistants, and Gwern has named a version of such systems guardian angels.
What's hard, though, is actually getting them made: while loyalty is a strong user preference, those creating very powerful AI systems and products won't necessarily design systems with loyalty coming first, precisely because they have their own interests in terms of engagement, profit, etc. (at minimum) and potentially more nefarious ones. There are potential routes through evals, certification processes, etc., and I think these are worth pursuing; but ultimately they can at best create some pressure that companies can evade to a large extent.
But an idea I came to recently is to separate, insofar as possible, the core capability and work done by an assistant from the fiduciary/loyal behavior part – and provide the second as a service. Such a "fiduciary overlay" would be more like an angel sitting on and reading over your shoulder (to awkwardly mix shoulder metaphors), and alerting you whenever something is awry. For example, when an AI system is being highly sycophantic, or has conflict of interest, or is being overconfident, or is possibly hallucinating, or is getting to a stupefying context-window length, etc., this little angel would let you know. And then you can decide what to do with that.
There's lots more to say here, but for now I'll just say that: the brand new Alloidal Foundation, in collaboration with Future of Life Foundation, is now trying to make this! Below is an RFP for doing some of the research work of figuring out, from an ongoing conversation, whether an AI system is falling into one of these failure modes, so that it can be flagged.
If this is up your alley, please take a look at the full RFP: https://allodial.org/rfp/ - proposals will be for $50K, with selections made by Oct. 2.
Questions are welcome in the comments, or write to hello@allodial.org.