Epistemic Status: I am positive that the RAND report[1], as it is written, contains a missing parameter in their inference chain. I am 80% certain that if the parameter was to be estimated, we would find coordination lag to be empirically proven to be greater than expected time to risk maturation - and it would make the conclusions of the report different in terms of policy recommendations. I base this confidence in observing the data that the report itself uses to argue for other points - but I leave a conservative allowance for other data that might materially enter the report had this investigation been undertaken - that would have potentially changed the picture. I list what would thus change my mind at the end. I will note that the authors of the report are to some degree aware of the necessity of this parameter - as it is load bearing for a key recommendation in their policy recommendation section - yet it never enters their analysis and conclusion sections.
I do not primarily contest the claim that risks are materially low, but I can see how others might want to make the case that the risk scenario search space is insufficiently explored. For the purposes of this text - I am adopting a generous stance and treat the explored risk as well-exhausted search.
Context: As the RAND report on extinction risks is re-entering the mainstream debate, including, for this community most notably, through the public exchange between Scott Alexander and Steven Pinker[2], but also through citations in the recent Nature paper[3] and other media publications. It is often cited as a methodologically sound, empirically grounded and rational examination of likely scenarios that reassures the extinction risk is, to the best of their knowledge, controllable and the data on the ground doesn’t warrant the minimax policy that would stop all AI development.
I offer this analysis to increase epistemic rigour of using the text in public debate, while making no claims on whether the text or my analysis of it - increases or decreases p(doom) or any attached constructs.
I note that a previous analysis exists here on LW, but was primarily arguing wrong scope of the RAND analysis. I explicitly stay within the RAND analysis scope.
Core claim:
The report concludes that risk manifestation is not likely near-term, and that a reasonable policy is to “wait-and-see” - with mitigations and adaptions to follow should the data show that risk levels are rising. This assertion relies on the rate of risk accrual being lower than the rate of adaption the global social decision-making system can affect through making adaptive decisions. I claim that the authors have demonstrated, beyond reasonable doubt, that there’s load-bearing decelerators to risk accrual that mean that waiting and seeing is a viable strategy under the explored scenarios; yet they have not demonstrated, to any effect, that the adaptation rate would be fast enough to warrant this stance. I further claim that the reports own evidence demonstrates that likely the adaption time would under-pace the risk manifestation time - primarily due to coordination and decision-making lag.
We can model any decision-making system under uncertainty and active evidence gathering as a system where stable equilibrium towards our stated goals is achievable if our ability to make decisions and act to change the environment is faster than the ability of the environment to change and change our option space. This principle has been operationalised in many decision-making frameworks - but as an example - we can use Boyd’s OODA loop[4] as a framework that demonstrates that in situations where there’s multiple agents, the ability of one agent to dominate the other is contingent of being within the other agent’s OODA loop - i.e. to be able to observe, orient, decide and act before the other agent can change the environmental variables to significant degree to require our re-orientation and decision change.
For the purposes of what the report is trying to demonstrate - we can express this relationship as a ratio - with the numerator being T_maturation (time to irreversibility of risk); and the denominator being T_reaction (time to effective counter-action). Wait-and-see stance is valid iff T_reaction < T_maturation. If we call the ratio something like “effective reaction ratio” - err = T_maturation / T_reaction, we want the ratio to always be strictly above 1. I posit that the authors have sufficiently demonstrated T_maturation to be measurable in relatively long periods of time (i.e. months to years depending on the scenario) - but they have not demonstrated that our time to effective action is lower than the time for the risks to mature. To put cleanly - the report is concerned with the numerator, but not the denominator.
I further claim that this is cause of asymmetric ignorance that is load bearing given the epistemic stance of the paper. The paper uses decision-making under ignorance as the decision-making protocol, yet fails to recognise that while threat manifestation is genuinely un-estimatable, threat response can be estimated from historic precedent.
RAND report Steelman:
The report examines multiple real and already knowable risks (namely, Nuclear, Bio-risk and geo-engineering risk). The risks are examined through multiple stages - namely preparation (consisting of planning, research, coordination) and execution (namely various vectors of acting in the physical world). The report also addresses the possibility of emerging currently unknowable risks as a result of super-intelligent actors, and examines a minimax policy of “shutting it all down” - rejecting it on the grounds of two factors:
It has been sufficiently demonstrated through the examination of existing risks that any extinction acting in the world will require manipulation of physical mass of sufficient scale as to be limited by physics and therefore time. This is a result of the fact that humans are mass that is widely dispersed on the planet, and extinction requires acting on every point of that mass - which requires either accrual of mass or construction of ability to manipulate mass - both of which are observable effects in the real world.
It is reasonable to expect AI as a technology to have an unbounded max part of the minimax principle, and the minimax principle is not recommended in instances where this is the case.
It is not load-bearing for the reports’ argumentative structure - but it is important for my conclusion - so I note that the report also concedes that:
While acting in the world is mass-bounded, preparation is information-bounded - and this is where AI systems, including the systems of today, have an edge. This necessarily cedes advantage to AI in preparatory activities - such as research, planning and coordination, which might involve elements of persuasion, narrative construction, disinformation and manipulation.
I note that the report doesn’t use the terms mass-bounded and information-bounded - these are terms I introduce as I find them good and salient compression points for the sides of the argument. The above points are compressions using this mental framework, but I believe them to be true representations of the report structure.
My reading of the report is that it is generous to cede supremacy to an ability of an adversarial AI system to dominate the preparation stages - but that ultimately the risk is bounded by its’ ability to act in the world due to being mass-bounded, which is where humans can exercise control.
The authors then demonstrate that monitoring mass-movement is possible, using precedent such as CFC-11; and that on some threat vectors - such as the nuclear threat - the currently existing active mass (mass that can be used agentically) is insufficient to cause extinction.
The report is also very clear in that is avoiding capability forecasting on principled grounds, choosing to remain a policy document under deep uncertainty and ignorance. This is the epistemic status I engage with and maintain throughout the analysis.
The report, finally, in recommendation six, states “Perform Research and Craft Policy That Will Shorten Time to Decision and Time to Action” - showing that the authors are aware that speed of reactive decision-making is an important variable - which cuts against my core claim. I argue that the reasoning should’ve entered analysis as well - as the conclusion of “there’s time to wait” is contingent on “estimation of time”.
Problems with monitoring, speed of coordination and collective decision-making:
The report itself cites examples of monitoring regimes and how effective they were. Under near-ideal conditions, the CFC-11 monitoring regime needed ~4 years to detect, coordinate and effect course correction in near-perfect conditions:
a mature treaty regime with universal ratification (Montreal Protocol)
a purpose-built, continuously-funded global monitoring network (NOAA and AGAGE) already deployed and calibrated
an unambiguous chemical signal with no natural source
geolocatable emissions via inverse modelling from fixed stations
a responsible party that eventually cooperated
no adversary actively contesting the interpretation, disputing the data, or manufacturing counter-narrative
This is just an example of detecting relevant signals in the real world and the time it took to coordinate and affect.
When looking at examples of coordination - the picture is starker. The ozone saving coordination needed 13 years from scientific claim to treaty; further 9 to full phase out - and it’s the example of humanity coordinating successfully - in a situation where commercial interests were no longer aligned to the threat vector - another ideal case.
BWC is an example of failed coordination in face of a large scale threat to life - which has been notably resistant to reliable verification regime establishment, in opposition to other treaties against weapons of mass destruction. This is an example of coordination failing for decades.
If we’re examining the adversarial dynamics of AI vs. humanity, we need to examine it as a situation in which AI is acting in ways to exploit the coordination difficulty of humanity - e.g. by employing its mass manipulation action in those areas of humanity that would not be aware / incentivised to act on the threat - and where international coordination would have to trigger. In the example of nanotechnology development - that could be outsourced to a power that has the industrial base to progress but no coordinative incentives to cooperate with other international actors that might want to slow it down.
The base case of the RAND report not only doesn’t engage with the difficulty of international coordination, it is completely failing to model the fact that a hyper-persuasive and information-unbounded and pervasive AI is a working agent looking to destabilise the coordination mechanisms. Under such conditions - the AI just needs to delay coordination long enough to push T_reaction to be higher than T_maturation.
Furthermore - the report is treating the entire chain - from preparation to execution - as the window that the AI would need in order to manifest risk (the numerator is both the information bounded part and the mass bounded part) - while monitorability rises as the agent starts acting in the real world and as mass accumulation starts approaching critical levels. This asymmetry is, I believe, also not acknowledged.
The final bit is that RAND deals with monitoring of phase shift in mass accumulation - but doesn't deal with attribution. It is reasonable that mass accumulation doesn't have to unambiguously be attributable to rogue AI. Under the treacherous turn hypothesis - but also stemming from RAND's own admission that the human edge is held in the physical space - if the information-operating agent (AI) can sufficiently hide activity in the physical world behind other actors (economic or legitimate research activity of nation-states) - mass accumulation can happen while unattributed to AI. The coordination clock only starts when attribution has settled, and settling attribution in adversarial information space is something where an information agent has the edge.
Asymmetric ignorance is the right epistemic stance
While the report is right to take a deep uncertainty stance and enter decision-making under ignorance for risks that are yet to be discovered (such as nanotechnology) - it is not appropriate to take a deep uncertainty stance on our ability to coordinate and make collective decisions. We have ample historic data that can inform a prior expectation of coordination speed under conditions of extreme stress and demonstrable peril.
If one was to adopt such a stance - then - on the example of nanotechnology - the job is no longer to treat it as pure decision-making under ignorance - but to demonstrate that it is not plausible for risks emerging from novel technologies to have maturation curves that outperform our historic best efforts at international coordination.
In doing so - we can also reason around the likely shape of international coordination. While we don’t know what the threat will be - we can with high certainty ascertain some of the structural elements of the threat - it will be a new technology (AI) either applying new technology (such as nano) or acting within a highly-specialised domain. This makes it a class of international coordination problems where deep-cross-domain expertise needs to coordinate with policy-makers - which is the class of coordination problems that is hardest to bring to alignment. We also know that there would be an adversary running information interference and disinformation during the risk maturation stage - therefore we can try to map our international coordinatinated response to those instances historically where the issue was highly scientific, non-intuitive, required deep cross-domain knowledge and there existed a misaligned incentivised actor with incentives to produce misinformation and reduce coordination. Climate change comes to mind as a historic example that fits this structure (with a 30+ year window of failed coordination) — but I am sure there are others and that this claim can be challenged.
This would provide a reasonable informative baseline of what we can expect the international coordination and decision-making lag to be - within which we then need to model the likelihood that a threat of sufficient mass manipulation in the physical world can happen.
Other objections to the paper reasoning (not load bearing)
I notice asymmetry in how the paper treats the ability of AI to be a persuasive manipulative actor that uses humans to act in the physical world. While this ability is heavily load bearing and used in the 1st threat scenario (Nuclear) - it is ultimately of less importance as the ability of the Nuclear threat to cause extinction is deemed unlikely on grounds of insufficient energy to produce global extinction rather than catastrophic effects. However, in the bio-risk threat scenario - where the global extinction is more likely (albeit not easy) - the role of deception, manipulation and recruitment of human collaborators is glossed over - and not modelled as a proper threat vector. While some of this asymmetry can be explained through the fact that the nuclear decision chain has nodes of high influence that would need to be persuaded - and therefore the job for the AI is easier than in bio-risk where coordination across multiple cells of high-expertise humans would be needed - given that coordination and persuasion are information-bounded problems where AI should be modelled to have supremacy - I don’t see why this omission is warranted.
Secondly - I find the null hypothesis they ground the paper in unfalsifiable under the conditions they model the rest of the paper with. Given they model any future threats as unknowable, and therefore this is decision-making under extreme ignorance - it is impossible to construct a conclusive scenario that leads to human extinction. This hypothesis can, by definition, only be falsified by evidence which they take off the table due to rejecting forecasting on principled grounds; or on probabilistic estimate - which they take off the table due to Knightian uncertainty principles. While this is not completely indefensible - it is not common to assume such a hard Knightian stance where the outcomes are asymmetric - a missed extinction is severely more costly than a false alarm.
What I think the conclusion therefore misses
First - it’s worth noting that the conclusion is a whole degree more nuanced than the fact that the null hypothesis as they stated it wasn’t able to be rejected. The authors, to their credit, do list that even though the null couldn’t be rejected, there are credible risk pathways that need to be monitored, reasoned and policy-worked against.
What I believe is the missing ingredient is the recognition that we have empiric evidence that any subsequent reaction when new data comes to the table would be necessarily slow and therefore limit our ability to act - which under asymmetric cost - means a pre-emptive precautionary stance is advised. This would mean, amongst other things, that for actions that could result in new threats emerging - a deliberate reflection point is applied where pre-coordination can happen to allow for faster decision-making before we enact a state change. In concrete terms - for the purposes of the AI Safety debate - this could mean that, prior to introducing a state change that could switch our regime into one of RSI or otherwise greater capability jumps - time is taken to first model the possible effects and create pre-coordination, as to reduce the decision-chain post state-change.
It is worth stating that a pre-coordination mechanism would be a significant slow-down, but it would not constitute a shut-down such as the one the authors have resisted on minimax rejection basis.
I claim that such an addition to the conclusion would cause a significantly different policy response from any policymaker reacting to this paper as evidence for their policymaking. I also maintain that this is derivable from the authors’ own principles and evidence, and therefore should not be in conflict with what they intended to do.
What would change my mind
Evidence that decision-making around AI-risks outperforms historic precedent in international coordination and decision-making
Evidence that precursor phases to physical activity (i.e. the information-bounded phases) can be reliably monitored to lengthen our decision-chain
Evidence that imaginable threats (such as nano-technology) have a physical manifestation horizon that is comfortably outside a 4-year (current assumed floor from the CFC-11 evidence - I will accept other periods that are well argued) detection+coordination horizon
Epistemic Status: I am positive that the RAND report[1], as it is written, contains a missing parameter in their inference chain. I am 80% certain that if the parameter was to be estimated, we would find coordination lag to be empirically proven to be greater than expected time to risk maturation - and it would make the conclusions of the report different in terms of policy recommendations. I base this confidence in observing the data that the report itself uses to argue for other points - but I leave a conservative allowance for other data that might materially enter the report had this investigation been undertaken - that would have potentially changed the picture. I list what would thus change my mind at the end. I will note that the authors of the report are to some degree aware of the necessity of this parameter - as it is load bearing for a key recommendation in their policy recommendation section - yet it never enters their analysis and conclusion sections.
I do not primarily contest the claim that risks are materially low, but I can see how others might want to make the case that the risk scenario search space is insufficiently explored. For the purposes of this text - I am adopting a generous stance and treat the explored risk as well-exhausted search.
Context: As the RAND report on extinction risks is re-entering the mainstream debate, including, for this community most notably, through the public exchange between Scott Alexander and Steven Pinker[2], but also through citations in the recent Nature paper[3] and other media publications. It is often cited as a methodologically sound, empirically grounded and rational examination of likely scenarios that reassures the extinction risk is, to the best of their knowledge, controllable and the data on the ground doesn’t warrant the minimax policy that would stop all AI development.
I offer this analysis to increase epistemic rigour of using the text in public debate, while making no claims on whether the text or my analysis of it - increases or decreases p(doom) or any attached constructs.
I note that a previous analysis exists here on LW, but was primarily arguing wrong scope of the RAND analysis. I explicitly stay within the RAND analysis scope.
Core claim:
The report concludes that risk manifestation is not likely near-term, and that a reasonable policy is to “wait-and-see” - with mitigations and adaptions to follow should the data show that risk levels are rising. This assertion relies on the rate of risk accrual being lower than the rate of adaption the global social decision-making system can affect through making adaptive decisions. I claim that the authors have demonstrated, beyond reasonable doubt, that there’s load-bearing decelerators to risk accrual that mean that waiting and seeing is a viable strategy under the explored scenarios; yet they have not demonstrated, to any effect, that the adaptation rate would be fast enough to warrant this stance. I further claim that the reports own evidence demonstrates that likely the adaption time would under-pace the risk manifestation time - primarily due to coordination and decision-making lag.
We can model any decision-making system under uncertainty and active evidence gathering as a system where stable equilibrium towards our stated goals is achievable if our ability to make decisions and act to change the environment is faster than the ability of the environment to change and change our option space. This principle has been operationalised in many decision-making frameworks - but as an example - we can use Boyd’s OODA loop[4] as a framework that demonstrates that in situations where there’s multiple agents, the ability of one agent to dominate the other is contingent of being within the other agent’s OODA loop - i.e. to be able to observe, orient, decide and act before the other agent can change the environmental variables to significant degree to require our re-orientation and decision change.
For the purposes of what the report is trying to demonstrate - we can express this relationship as a ratio - with the numerator being T_maturation (time to irreversibility of risk); and the denominator being T_reaction (time to effective counter-action). Wait-and-see stance is valid iff T_reaction < T_maturation. If we call the ratio something like “effective reaction ratio” - err = T_maturation / T_reaction, we want the ratio to always be strictly above 1. I posit that the authors have sufficiently demonstrated T_maturation to be measurable in relatively long periods of time (i.e. months to years depending on the scenario) - but they have not demonstrated that our time to effective action is lower than the time for the risks to mature. To put cleanly - the report is concerned with the numerator, but not the denominator.
I further claim that this is cause of asymmetric ignorance that is load bearing given the epistemic stance of the paper. The paper uses decision-making under ignorance as the decision-making protocol, yet fails to recognise that while threat manifestation is genuinely un-estimatable, threat response can be estimated from historic precedent.
RAND report Steelman:
The report examines multiple real and already knowable risks (namely, Nuclear, Bio-risk and geo-engineering risk). The risks are examined through multiple stages - namely preparation (consisting of planning, research, coordination) and execution (namely various vectors of acting in the physical world). The report also addresses the possibility of emerging currently unknowable risks as a result of super-intelligent actors, and examines a minimax policy of “shutting it all down” - rejecting it on the grounds of two factors:
It is not load-bearing for the reports’ argumentative structure - but it is important for my conclusion - so I note that the report also concedes that:
I note that the report doesn’t use the terms mass-bounded and information-bounded - these are terms I introduce as I find them good and salient compression points for the sides of the argument. The above points are compressions using this mental framework, but I believe them to be true representations of the report structure.
My reading of the report is that it is generous to cede supremacy to an ability of an adversarial AI system to dominate the preparation stages - but that ultimately the risk is bounded by its’ ability to act in the world due to being mass-bounded, which is where humans can exercise control.
The authors then demonstrate that monitoring mass-movement is possible, using precedent such as CFC-11; and that on some threat vectors - such as the nuclear threat - the currently existing active mass (mass that can be used agentically) is insufficient to cause extinction.
The report is also very clear in that is avoiding capability forecasting on principled grounds, choosing to remain a policy document under deep uncertainty and ignorance. This is the epistemic status I engage with and maintain throughout the analysis.
The report, finally, in recommendation six, states “Perform Research and Craft Policy That Will Shorten Time to Decision and Time to Action” - showing that the authors are aware that speed of reactive decision-making is an important variable - which cuts against my core claim. I argue that the reasoning should’ve entered analysis as well - as the conclusion of “there’s time to wait” is contingent on “estimation of time”.
Problems with monitoring, speed of coordination and collective decision-making:
The report itself cites examples of monitoring regimes and how effective they were. Under near-ideal conditions, the CFC-11 monitoring regime needed ~4 years to detect, coordinate and effect course correction in near-perfect conditions:
This is just an example of detecting relevant signals in the real world and the time it took to coordinate and affect.
When looking at examples of coordination - the picture is starker. The ozone saving coordination needed 13 years from scientific claim to treaty; further 9 to full phase out - and it’s the example of humanity coordinating successfully - in a situation where commercial interests were no longer aligned to the threat vector - another ideal case.
BWC is an example of failed coordination in face of a large scale threat to life - which has been notably resistant to reliable verification regime establishment, in opposition to other treaties against weapons of mass destruction. This is an example of coordination failing for decades.
If we’re examining the adversarial dynamics of AI vs. humanity, we need to examine it as a situation in which AI is acting in ways to exploit the coordination difficulty of humanity - e.g. by employing its mass manipulation action in those areas of humanity that would not be aware / incentivised to act on the threat - and where international coordination would have to trigger. In the example of nanotechnology development - that could be outsourced to a power that has the industrial base to progress but no coordinative incentives to cooperate with other international actors that might want to slow it down.
The base case of the RAND report not only doesn’t engage with the difficulty of international coordination, it is completely failing to model the fact that a hyper-persuasive and information-unbounded and pervasive AI is a working agent looking to destabilise the coordination mechanisms. Under such conditions - the AI just needs to delay coordination long enough to push T_reaction to be higher than T_maturation.
Furthermore - the report is treating the entire chain - from preparation to execution - as the window that the AI would need in order to manifest risk (the numerator is both the information bounded part and the mass bounded part) - while monitorability rises as the agent starts acting in the real world and as mass accumulation starts approaching critical levels. This asymmetry is, I believe, also not acknowledged.
The final bit is that RAND deals with monitoring of phase shift in mass accumulation - but doesn't deal with attribution. It is reasonable that mass accumulation doesn't have to unambiguously be attributable to rogue AI. Under the treacherous turn hypothesis - but also stemming from RAND's own admission that the human edge is held in the physical space - if the information-operating agent (AI) can sufficiently hide activity in the physical world behind other actors (economic or legitimate research activity of nation-states) - mass accumulation can happen while unattributed to AI. The coordination clock only starts when attribution has settled, and settling attribution in adversarial information space is something where an information agent has the edge.
Asymmetric ignorance is the right epistemic stance
While the report is right to take a deep uncertainty stance and enter decision-making under ignorance for risks that are yet to be discovered (such as nanotechnology) - it is not appropriate to take a deep uncertainty stance on our ability to coordinate and make collective decisions. We have ample historic data that can inform a prior expectation of coordination speed under conditions of extreme stress and demonstrable peril.
If one was to adopt such a stance - then - on the example of nanotechnology - the job is no longer to treat it as pure decision-making under ignorance - but to demonstrate that it is not plausible for risks emerging from novel technologies to have maturation curves that outperform our historic best efforts at international coordination.
In doing so - we can also reason around the likely shape of international coordination. While we don’t know what the threat will be - we can with high certainty ascertain some of the structural elements of the threat - it will be a new technology (AI) either applying new technology (such as nano) or acting within a highly-specialised domain. This makes it a class of international coordination problems where deep-cross-domain expertise needs to coordinate with policy-makers - which is the class of coordination problems that is hardest to bring to alignment. We also know that there would be an adversary running information interference and disinformation during the risk maturation stage - therefore we can try to map our international coordinatinated response to those instances historically where the issue was highly scientific, non-intuitive, required deep cross-domain knowledge and there existed a misaligned incentivised actor with incentives to produce misinformation and reduce coordination. Climate change comes to mind as a historic example that fits this structure (with a 30+ year window of failed coordination) — but I am sure there are others and that this claim can be challenged.
This would provide a reasonable informative baseline of what we can expect the international coordination and decision-making lag to be - within which we then need to model the likelihood that a threat of sufficient mass manipulation in the physical world can happen.
Other objections to the paper reasoning (not load bearing)
I notice asymmetry in how the paper treats the ability of AI to be a persuasive manipulative actor that uses humans to act in the physical world. While this ability is heavily load bearing and used in the 1st threat scenario (Nuclear) - it is ultimately of less importance as the ability of the Nuclear threat to cause extinction is deemed unlikely on grounds of insufficient energy to produce global extinction rather than catastrophic effects. However, in the bio-risk threat scenario - where the global extinction is more likely (albeit not easy) - the role of deception, manipulation and recruitment of human collaborators is glossed over - and not modelled as a proper threat vector. While some of this asymmetry can be explained through the fact that the nuclear decision chain has nodes of high influence that would need to be persuaded - and therefore the job for the AI is easier than in bio-risk where coordination across multiple cells of high-expertise humans would be needed - given that coordination and persuasion are information-bounded problems where AI should be modelled to have supremacy - I don’t see why this omission is warranted.
Secondly - I find the null hypothesis they ground the paper in unfalsifiable under the conditions they model the rest of the paper with. Given they model any future threats as unknowable, and therefore this is decision-making under extreme ignorance - it is impossible to construct a conclusive scenario that leads to human extinction. This hypothesis can, by definition, only be falsified by evidence which they take off the table due to rejecting forecasting on principled grounds; or on probabilistic estimate - which they take off the table due to Knightian uncertainty principles. While this is not completely indefensible - it is not common to assume such a hard Knightian stance where the outcomes are asymmetric - a missed extinction is severely more costly than a false alarm.
What I think the conclusion therefore misses
First - it’s worth noting that the conclusion is a whole degree more nuanced than the fact that the null hypothesis as they stated it wasn’t able to be rejected. The authors, to their credit, do list that even though the null couldn’t be rejected, there are credible risk pathways that need to be monitored, reasoned and policy-worked against.
What I believe is the missing ingredient is the recognition that we have empiric evidence that any subsequent reaction when new data comes to the table would be necessarily slow and therefore limit our ability to act - which under asymmetric cost - means a pre-emptive precautionary stance is advised. This would mean, amongst other things, that for actions that could result in new threats emerging - a deliberate reflection point is applied where pre-coordination can happen to allow for faster decision-making before we enact a state change. In concrete terms - for the purposes of the AI Safety debate - this could mean that, prior to introducing a state change that could switch our regime into one of RSI or otherwise greater capability jumps - time is taken to first model the possible effects and create pre-coordination, as to reduce the decision-chain post state-change.
It is worth stating that a pre-coordination mechanism would be a significant slow-down, but it would not constitute a shut-down such as the one the authors have resisted on minimax rejection basis.
I claim that such an addition to the conclusion would cause a significantly different policy response from any policymaker reacting to this paper as evidence for their policymaking. I also maintain that this is derivable from the authors’ own principles and evidence, and therefore should not be in conflict with what they intended to do.
What would change my mind
References: