Let's say your company solves value generalization. And meanwhile, the AI industry evolves in such a way that the strongest AIs are controlled by a handful of very rich and powerful people. You sell them your research, they use it to align AI to themselves, and we end up in a world with overlords. To me this would be a bad outcome. Do you think it would be a good outcome? Or do you intend to use your research in some other way?
That's why part of the plan is to sell the models (or use of the models), not to sell the research https://www.lesswrong.com/s/EYgCdcxsn73fKWWra/p/uMKGaEKRDpoqnZyBh
And that's the main reason that I'm considering the commercial path in the first place, to get some control over the use of the technology.
Dividing the possible world into "value generalisation has strong capability increases": then I can get powerful models with much less investment (today's models are ridiculously overtrained to compensate for their poor generalisation). And "value generalisation doesn't have strong capability increases": then the risk is lower.
The bad spot would be "value generalisation has strong capability increases that only large established companies can take advantage of".
Hmm, to me your reply makes the idea even more surprising and dangerous. Basically your plan is to train better models than OpenAI/Anthropic/etc. Either this won't work, or it will work and attract all the world's money and power. There is no "structure to prevent investor control or excessive pressure from profit motives" that can withstand this.
I doubt there's no such structure, but the amount of load it must bear is very significant. it would be a remarkable achievement of organization design.
Recall that this would be under the assumption that: "it is deemed safe (by us and the AI ethics board) to forge ahead commercially (or at least safer than not doing so)". If that assumption fails, then we'd have shut down or pivoted to immediate profit instead of doing this.
So if we go ahead commercially, we don't have to resist excessive pressure - we just need to be commercially successful within the standard parameters. We also don't need to be strictly better than OpenAI or Anthropic - just successful in specific sub-markets or sub-uses. One option is to build on that success with a possible licensing system - "you can use our technology, for a reasonable fee, as long as it's also using value-aligned models". I'm hoping that pre-aligned models will be so successful in their own niche that they set the design for that these models should be.
How would you prevent your researchers from being hired at, say, 10-figure salaries, realizing much more of their expected future profits that your board denies them, and with at least as much certainty as the short-term cash you offer, and getting to pursue their cutting-edge research?
Nothing is ever guaranteed, but we would recruit moral people. One possible approach would be to have contracts that are much more restrictive (in terms of IP and working on the ideas at a rival) in the case of ethics board blocking, rather than otherwise. That's a good idea; thanks for prompting the thought.
Well, California rather famously treats precisely the kinds of contract provisions you're envisioning as void, so you'll first need to be out-of-state and sue for federal injunctions enforcing your terms against your "moral people" if they quit, who would need to be exceedingly naïve to sign on after all the unusual-for-the-field and obviously adversarial measures they'd see you taking just to be able to abuse them later, or lacking in options.
So, careful legal analysis before doing any of this, got it :-)
Thanks for the feedback. Feasibility conversations at the beginning are always useful, even if legal experts would catch the issues later on.
Could you explain the alternative? Suppose that MIRI/Prealign/etc solves value generalization and is forced to choose between OpenBrain creating Agent-5 and OpenBrain creating Safer-4. Then any reasonable person would choose Safer-4 over Agent-5... unless there emerges a way to have OB align Safer-4 to the humans in general (e.g. by using Chinese auditors from AI-2040's Plan A).
I don't trust myself as sole dictator (my ethics might hold, but my judgement might be impaired), so the structures will have to be able to override me if I go unbenevolent.
Because it's not a case of virtue with a clear line ("never accept a bribe") or virtue with an unclear line ("at what point does networking turn into nepotism"), but a case of careful judgement in the presence of strong confounding incentives. In that case, outside opinion is invaluable, but needs teeth.
Though thank you for the implied trust :-)
OpenAI was founded as a non-profit with a clear commitment to avoiding AI dangers and a unique structure to prevent investor control or excessive pressure from profit motives.
This failed.
I'm considering launching a commercial venture (possibly named "Prealign") to solve value generalisation and contribute to AI alignment, while avoiding anything that increases AI capabilities relative to alignment and values.
I want a strong mission lock to keep the company ethical and an AI ethics board with real teeth and power to enforce that.
But seeing past failures, it's obvious we have to be careful here. So I'm asking for ideas to make sure that "Prealign" or whatever the organisation ends up being called, actually does stay aligned.
I've been mainly thinking along conventional structural lines - maybe the ethics board should own the company, different share classes, incorporating in locations where a mission lock in the company's founding documents is respected and enforceable, choosing the right corporate design, etc.
Any specific idea of that sort is of value, as are unconventional or original ideas as well. I'll just share one unusual design thought I had: the pivot to profit.
When things look dangerous, pivot to immediate profit
The OpenAI situation highlighted that it's not just a question of the leadership holding out against immense financial and social pressures; you have to consider the employees and partners as well. If the company holds out but the employees jump ship[1] to a deep-pocketed partner or rival, carrying their knowledge with them, little will have been achieved.
Now, in the normal course of operation, a company focused on deep research will turn up lots of designs, results, and customer needs that are valuable but not worth working on, because the deep research is the focus and priority.
So I suggest that if the deep research starts looking dangerous, the ethics board has the power to impound it and announce a pivot to immediate profit. The cutting edge research is taken off the table; and the company has to turn instead to monetising all those previous designs, results, and connections.
So instead of fighting profit head-on, it would go with it. Trading future profits and future immense risks for much more certain short-term cash. Now, people might certainly be disappointed (I know I would be) but it shouldn't provoke the all out fight that shutting down the company would.
How does that idea sound?
Of course we'd want to only employ moral and dedicated people; but here's an omnipresent self-serving bias in everyone. It's easy to convince yourself (yes, you, yourself, and also me, myself), in ambiguous situations, that things are "not quite so bad" and that the AI ethics board are being unreasonable or overcautious. If the design depends on everyone being saintly and resistant to pressure, it will fail.