The point. We all are confidently overlooking the root cause of why the alignment problem exists and I claim to have found a cluster of reasons but I cannot convince the right (amount) of people to take them seriously to make any progress on it. It relates to ML-involving technology at many levels of specificity and many points of the supply chain that enables it. (Also, please suspend your expectation that to meaningfully address a fundamental problem as this, using highly technical language or invoking high-context concepts is absolutely necessary, in fact I shall address this in particular On why I am admitting defeat in trying to articulate this.)
The short, poetic-intended version
With that, recall how the paperclip maximizer story used to be a go-to metaphor of how devastatingly risky it could be to unleash an ASI (if super-humanity ever was crucial to it in the first place) for 3 reasons at minimum:
it gave compelling, tangible picture of how far existentially threatening impact could go, to let itself capture the imagination of a wide public;
it was careful to be anchored to firm, "super-paradigmatic"[1] ideas like optimization and reward misspecification, to make itself last;
it proposed a problem in a way that, if not directly inspiring to think about it (even in technical terms), was at least suggestive of optimism, to make the audience believe that humanity is rather likely to solve it or manage it and live happily ever after.
Squinting at the paperclip maximizer, what does it let you see? As for me:
we entertain the possibility of engineering something that is good at making sure it will not stop operating; among its ingredients:
a polestar-like incentive (towards which ever-progressing steps can be taken without ever reaching it),
the possibility of some steps or activities being instrumental in the sense that all downstream approximation of the target can be believed to be more rewarding and / or efficient by some metric.
And of course, other details as well, which may or may not be necessary for that outcome, but the outcome is a universe consisting entirely of paperclips -- something meant to be absurdly ironic.
For what can this be a fitting metaphor?
Observe that it implies the belief in -- perhaps also a desire for -- a collective self-imposed ontological shift: the "we", who succeeds in engineering the form of intelligence that we are looking for here "correctly", is expecting that that very act is so rewarding as to never want to revert it. On the contrary, such success is an achievement that is instrumental to the prosperity of the collective, long term, so why would we?
An aside on mythologies
There are 2 assumptions that this takes vitally seriously, rooted in elements of culture, one more contemporary and another more ancient (perhaps archaic to some readers), respectively.
Firstly, the tenet that additional pieces of knowledge, behavioral traits, mechanisms etc. when viewed in a historical context consistently tend to improve the prosperity of the population. This is not really logically defendable if we adopt the position that ontologies shift throughout history and become increasingly incomparable. It does completely make sense as a belief that implies that the current living experience is not just what we can best relate to, but also the one we would be approximately the most grateful to have among all possibilities from history. It is a valuable mechanism to feel good about being grounded in physical reality.
The other is an important motif in the Judeo-Christian creation myth: any step of the creation starts with an act of separation that is contingent and not inherently permanent or even stable, and it ends with the meta-level observation that this is good. Accepting this creation myth not even at the level of making it part of a belief system, but just as a semantically correct and coherent text, asks for the context that the path from seeing that this is good to this being a permanent, even foundational feature of the universe is trivial. More explicitly: good should be what looks evidently, conspicuously in front of us and also that certitude and it being the basis of goodness should persist.
Observe furthermore that an incentive can be made polestar-like trivially as long as it involves an ideal that can be specified in ever-increasing precision and the deviation from that can be measured ever-increasingly precisely. I claim this trying not to restrict the definition of "precision in specification and measurement" arbitrarily at all. The best I can put it is: anything counts as higher precision if characterizing two things by the same scheme it (the more precise thing) overwhelmingly tends to take more information than the other, whatever the scheme.
Then note also that it seems intuitive that an incentive, like the incentives we humans like to follow, be best described just by some metric rather than criteria to meet. If we state something like "we just want to live healthier" but we can not quite delineate what that looks like, we can still likely come up with something to measure and by which to score our health, so obviously healthier means achieving a higher score. Wow -- this immediately even supplies a formulation of our goal that could be operationalized -- all we need is to play the game of pushing the score up and we should expect to live healthier! I think I am conveying here no more than the Bellman Q theorem.
Putting all of this together: what ticks boxes 0, 1 and 2 above is the very coordination that we see ourselves enact in the pursuit of ever-improving prosperity. I think we quite universally believe that
that coordination to work immediately has to seek self-preservation;
it also necessarily involves gathering information about the world, incl. ways in which we might actually prosper even more;
for trusting the coordination, moreover to engineer it, we need to measure its success and its own health;
additionally part of its purpose or its very existence is driving that metric up simultaneously.
Conclusion
Something may seem conspicuously imprecise with the above analogy: I have not yet mentioned any way in which the paperclip itself is an appropriate element. I shall do that now by answering "What is the product, the creation of which we are assigning to superhuman coordination?"
I find the paperclip to be the symbol for something that is both directly conducive to reward and the end product of instrumental action sequences. If the subject, in our case, superhuman coordination has begun to drive the "health metric", which is designed to accelerate all of (a) some object-level production, (b) instrument design and (c) meta-level strategy improvement (i. e. learning), then an indicator that this has happened is a signal that is progressively increasing and its increase is accelerating roughly proportionally to its temporary value, i. e. and exponential signal. More simply I want to say that the subject having recognized and acting out an abstract, widely applicable self-improvement policy, it should robustly produce some exponential signal, and that is what the paperclip stands for. Conversely, the symbol could also be the moles of matter consumed in (paperclip) production (which in case paperclips are standard size is just proportional to paperclips produced).
A natural next question should be "Which signs can all this point to that expose that some form of superhuman coordination is on track both to preserve itself and to dangerously maintain the consumption of something essential to (human) life?" And personally I want to follow that up with some that were also instructive to ask when I first encountered the paperclip story
Are there ways it can possibly be stopped?
Or slowed down?
Can we know what it takes to even slow it down?
If we know, can we do it?
etc.
Progress status
If you're reading this, you likely know that the text is left as a draft, to put it bluntly, because for all my thinking and research my answer to "If we know what it takes to slow down a current paperclip-type superhuman coordination, then can we do that?" is a convinced no. And that made me pretty devastated and depressed and the reason I still decided to share this with you is that I've some hope left that you care and want to change my mind.
The (slightly) more detailed and quantified version
Take statements that roughly have the shape "Heavily based on empirics, and perhaps impossible to ever cleanly deduce from theoretical first principles, yet it seems to robustly hold that sociometrical or economical signal is exponential (in , i. e. time)." Or "Welfare is characterized by a robust >1 proportional growth over a sliding window, yielding the expectation of a corresponding signal." My favorite examples are
Moore's law
for its striking polymorphism: it can be stated (with varying accuracy) for transistor density, per transistor chip price, FLOPS, global annual sale billings, and the list could go on with more combinations of the same underlying quantities, many of which I probably can't even imagine. Some that I can and should check: annual number of dies produced and sold, total annual chip number or FLOPS capacity increase via obsolescence and new acquisitions at big tech service and infrastructure providers, total annual turnover at enterprises in the supply chains of frontier entities (which produce anything from modules in ICT architecture, IC's, packaged chips, masks, wafers etc. to instrumentation, R & D, quality control, logistics) etc.
Even facing pessimism over the domain of the law due to the quantum tunneling limit on line width, it seems to "find a way", e. g. with Huawei's (now not very) recent announcement of logic-folding (stacked and bonded circuit layers), or with translations to different paradigms such as the quantum Moore's law.
I'm sure many people consider AI capability benchmark charts such as Time Horizon by METR to be in the spirit of the law even though if so, it has strayed significantly from the initial framing. There seems to be an implicit understanding that the spirit is not limited to transistor crafting but quantifies IT progress in general.
GDP
for its ironic function as a welfare metric: to my knowledge even at the conception of its use as such the reasoning behind it was primarily the mere practicality of its simple measurability and verifiability. That it is causally interconnected with more tangible, mundane experiences of welfare of nationals, is contestable at best.
Nonetheless, multiple natural factors must contribute to steady GDP growth even to maintain a constant chance of providing basic services to citizens. Most obviously a positive inflation rate is needed to maintain any currency's animating effect on economy. Even if we could somehow separate domestically and transnationally consumed components and adjust them for the respective inflation rates, we should expect to find steady growth reflecting at least that most populations are growing so even at constant per capita consumption total consumed wealth follows that trend (considering both immediate and estimated discounted future consumption). Of course, that is becoming complicated to compute by itself and is compounded with the fact that population growth is wildly variant and on the cusp of reversing in some of the more populous countries, too. But then if we consider a stagnant population and globally predictable inflation, we still have to take into account that the future window and portfolio of risks for which we want to prepare remediation can always be extended justifiably by the rationale of improving survival chances.
At least in 2 ways GDP growth still is instrumental to increasing living standards. On the one hand it provides a basis of comparison for nation economies: an aspect in which similarity can be assessed. One may learn important lessons from comparing economies undergoing periods of similar expansion either in living standards or GDP and examining the direct influence of social policy on how economies realize the affordances of pure productivity growth. On the other hand, states that aim at higher living standards that they perceive in other economies, and to that end implement some strategy, can use it as feedback on their effectiveness. This has arguably worked for several country development stories of the 20th century.
Now I want to look at some quantity like semiconductor market size over GDP per capita, to express the share of semiconductor industrial production in each person's contribution to the economy. More colloquially, I'd like to know the expected proportion of someone's added value coming from processing natural resources into chip substrates, and more concisely, estimate how much citizens are dedicated to wafer manufacturing on average. (Or, somewhat bluntly, our relative apparent need for silicon.) With all this elaboration I'm trying to be precise about my subject because I'm not sure if the quantity I end up investigating truly is the fitting and transparent choice, but more importantly, if I can convincingly argue that it is, both of which are crucial to make my point.
The presentation dilemma
The choice of quantity must be unbiased, to no extent cherry-picked to support my point, but honestly reflecting what I'm concerned with (again, quantifying our collective obsession with wafer production); and then it has to be meticulously researched, with the same intent. With that I haven't gone through, and it's worth mentioning why: as far as I can see, disentangling concurrent effect mechanisms that influence macroeconomic data series (especially ones that I have legal and / or physical access to) is rather complicated. Definitely so relative to the simplicity of the mathematics in the argument ("physicist's first approximation") that I wish to make. In addition, the argument is politically not neutral and nor is my personal affective attitude. Hence my motivation to see if I can make hard-science-supported, robust claims to argue for actionable, impactful interventions or shifts in culture. All that fits the type of manipulative propaganda all too well for my conscience.
To make matters worse I even strongly hold on to an intuition of what the shape of the simple mathematical result to derive and present should be.
The observation
I'm expecting the polymorphism, the wide-spread techno-optimism (or techno-fatalism?) of Moore's law and eagerness to extend its scope to various branches of IT and its supplying industry to manifest in several economic metrics growing in exponential order over time. Including ones that describe raw material extraction in particular and externalities or risks thereof in general. Meanwhile, GDP per capita (irrespective of subjective correlation to quality of life) stipulated to grow in exponential order and that being widely attempted to attain leads me to trust that such growth shall be observable long-term.
Then consider 2 corresponding signals, and respectively, which can be chosen so that their ratio adequately describes the quantitative answer to my question at the top of this section.
Around this year, indexed 0,
The logarithm of the ratio , let the ratio be , i. e. 's have the same linear term in their Taylor expansions. (Moreover, differences in higher order terms can only exist if the ratio is approximately constant and has a temporary Gaussian deviation from its current value.)
This tells us that unless we can tightly control per capita resource consumption, it exponentially grows relative to production (that is, conditionally on exponential scales of the quantities to begin with).
For an th order polynomial per capita extraction in but in the transient even that may generate an exponentially growing ratio, only asymptotically works.
That leaves us with very few options to avoid the blow-up of resource extraction to the point of rapid depletion and an eventual major economic collapse; depending on hard policy principles we can
relax control tightness: allow polynomial order consumption rates as control targets and take care of transients when shifting actual rates;
practically fix a per capita consumption rate (to GDP?);
forbid both exponential market expansion and exponential GDP per capita increase.
Which principles allow each?
Accepting the risk of poorly designed or executed market control (e. g. due to negligence) and allocate efforts to the collateral damage -- current regime, if grossly negligent or possibly non-existent market interference counts.
Stipulation of commonly accepted ideology in strong interventions on markets -- risk of ever-too-socialist for the US.
Use political leverage to control both rates -- growth taxation.
A proposal draft
"Taxation" in the name is not meant entirely literally: it has more to do with 2 notions referred to by the word[2]
the burden of corrective, reconstructive, precautionary or otherwise remedial action collateral to some primary operation of an enterprise;
a constraining mechanism to directly impose on markets that is capable of controlling price levels.
I want to be upfront that I find it easier to conceive of my proposal as that of a concrete form of tax, however establishing the very machinery needed to enforce a tax is in tension (if not direct contradiction) with the intent of the proposed instrument. (I shall address this shortly, and most directly in The intent.) Although this section looks like a report of brainstorming the proposal, its relevance to this text is not its content but being an example of addressing a problem type. I did spend significant time and effort on coming up with details that I'll eagerly share and debate elsewhere if anyone signals interest.
Fundamentals
DISCLAIMER: this iteration of the draft contains excerpts from conversations with and artifacts by Claude, of which the machine-generated parts I'm going to notate using quotation blocks.
The effect mechanism: progressively with the adoption of upgraded, supposedly improved versions of hardware or software by participants of an economy they must contribute from their capital to the management of possible and eventual damages and externalities occurring at any point of the involved supply chain, e. g. by financing a remediation fund.
The operative domain: transactions in which a user first receives access to an upgraded piece of technology that they had a prior version of; the tax can be priced into the transaction value but it is negotiated on a per-customer basis.
Basis: the customer's activity, granting that their own production or service, social responsibility, criminal record, feedback for the developer etc. may warrant a reduced bar of entry (tax deduction) for the upgrade; however this must be arbitrated by an impartial body, e. g. an NGO that neither the producer, nor the vendor and nor the customer have common or conflicting interests with.
Particular addenda
to simplify screening the subject, the tax amount can be progressively recalculated during the process, e. g.
proof of identity can be enough for a default but possibly prohibitively high base
by setting up increasingly restrictive claims about their intended use they can attain significant reduction
the amount of reduction is tied to the transparency and simplicity of auditing the commitments
in the transfer contract some form of audit must be specified that is minimally required to follow through with, and the arbitration body may unilaterally impose certain types of audit (which can be prescribed in subsequent legislation)
a hierarchical system of baseline audits has to be ratified, which are more general and wider economic sector-targeted at a higher hierarchy level and more specific for more unique products; these are to be included in the default tax agreement and executed as random checks
Tax applies at all supply-chain stages; each actor pays for direct externality of immediate upstream supplier only; cascading effect [of externality re-estimation] is indirect (triggered on-request at downstream customer's next upgrade, with favorable recalculation if upstream decreased externalities).
Upgrade[-specific audit criteria] determination case-by-case with reusable precedent; definition spans product, service, or market position, not technology alone; first-time entrants can claim "significant shift" [in their operation thanks to upgrading] (new entry) with NGO expert board assessment, [yielding a] zero-tax contract with auditing obligations.
Expert board: reassembled per case from standing credentialed registry, sector-technical expertise required, random selection.
Notification: immediate upon report creation, can occur at multiple points in 3-stage arbitration[ namely the breakdown of] NGO arbitration [into a] prospective (pre-clearing every buyer), real-time at point of sale, [and] retrospective/audited-after [stage].
Developers see aggregate trends [of their peers' tax contract commitments and compliance] + outlier flagging; consumers notified immediately; reports not public; outlier monitoring active (NGO + supplier); outlier reporting rewarded by severity/accuracy-based tax reduction tiering.
Tax visibility: public transparency on item description (externality category and context, format similar to US emissions reporting); downstream marketing not regulated.
Central registry [for expert board assembly]: government-maintained, cross-border capable, state surveillance accepted.
Regressive [tax burden compounding] offset: embedded in tax-rate variance tied to externality/efficiency basis, see Fundamentals.
Audit hierarchy: international standards set legal-category minimums; national jurisdictions elaborate; empirical sector selection.
First-time market entrants subject to tax unless significant shift claimed; transparency required to distinguish hype-driven positioning from genuine supply-chain entry; overhype risks audit failure or customer protection law violation.
Patches
This is a selection of ambiguous parts of the proposal that need precision and references to existing legal instruments that inspire in how to cover them (without claiming completeness).
Closest analogues. The PSLRA safe harbor for forward-looking statements (15 U.S.C. §78u-5) and the judicial "puffery"/"bespeaks caution" doctrines, plus Rule 10b-5 anti-fraud enforcement. Courts treat vague optimism ("transparently aspirational statements... mere corporate puffery") as non-actionable, but a concrete misrepresentation embedded in a narrative loses protection (e.g., In re Quality Systems, 9th Cir. 2017, on "mixed statements").
Co-occurring feature to add. The safe harbor is paired with a "meaningful cautionary language" requirement — protection only attaches if the statement is identified as forward-looking and accompanied by specific, non-boilerplate risk warnings. Your patch could grant audit immunity for narrative claims only when paired with filed, specific cautionary disclosures, exactly as the PSLRA conditions its safe harbor.
Closest analogues.FICO credit scoring is the canonical example of fixed per-category weight caps. Per Fair Isaac's official myFICO site, the score groups data into five categories: "payment history (35%), amounts owed (30%), length of credit history (15%), new credit (10%) and credit mix (10%)" — each category capped by a fixed maximum weight so no single factor dominates. In tax law, the General Business Credit aggregate cap (IRC §38(c)) bundles 30+ component credits and limits their combined total relative to tax liability. The US Sentencing Guidelines provide a structured-discretion analogue (fixed-size adjustments like the 2–3 level acceptance-of-responsibility reduction under §3E1.1).
Co-occurring feature to add. The Sentencing Guidelines pair structured caps with a departure/variance mechanism allowing reasoned deviation in extraordinary cases (with appellate review). Your independently-capped taxonomy could add a documented, reviewable "departure" channel for atypical cases that the caps fail to capture.
Patch 7 — Risk-based stratified random audit sampling with guaranteed minimum floor
Closest analogues. The IRS Discriminant Function (DIF) scoring system selects returns by statistical deviation from peer norms, stratified by taxpayer category; the IRS also runs random National Research Program (NRP) audits to calibrate the model — a built-in minimum-coverage floor across strata. Customs risk management under the WCO SAFE Framework is a parallel for trade.
Co-occurring feature to add. The IRS pairs scored selection with a dedicated random baseline program (NRP) that audits a floor of returns regardless of score, both to gather calibration data and to deter gaming. Your guaranteed audit floor is precisely this; consider explicitly designating the floor sample as also serving model-recalibration, as the IRS does.
Patch 9 — Tiered credentialed expert roster with random panels guaranteeing a senior member
Closest analogues.ICSID arbitration (World Bank): tribunals must be a sole arbitrator or any uneven number (default three), drawn from a Panel of Arbitrators, with majority-nationality-balance rules and a default appointment formula under Art. 37(2)(b). FDA advisory committees (21 CFR Part 14): committees require subject-matter expertise, quorum rules, and can add temporary voting members when needed expertise is absent.
Co-occurring feature to add. ICSID pairs panel composition with a disqualification/challenge procedure (arbitrators can be challenged for lack of independence, with a defined decision-maker). Your random-panel mechanism should add a structured conflict-of-interest screening and challenge process — FDA likewise screens every member for financial conflicts before each meeting.
Patch 10 — Lagged public release of aggregate disclosure data
Closest analogues.SEC Form 13F (15 U.S.C. §78m(f)): institutional managers (those with $100M+ in Section 13(f) securities) report holdings within 45 days after quarter-end, creating a minimum 45-day public lag explicitly designed to balance transparency against front-running/competitive harm. The EPA Toxics Release Inventory has an analogous reporting lag (prior-year data filed by July 1).
Co-occurring feature to add. Form 13F pairs the standard lag with a confidential-treatment request mechanism allowing additional delay (a quarter or more) where premature disclosure would harm an ongoing position. Your quarterly-lagged release could add a petition process for case-by-case extended confidentiality for especially sensitive competitive data.
Patch 14 — Earmarked sub-fund for the most vulnerable supply-chain tier
Closest analogues. The EU Just Transition Fund (Reg. (EU) 2021/1056), the grant pillar of the Just Transition Mechanism, earmarks money specifically for the territories and workers most negatively affected by decarbonization (coal/peat/oil-shale extraction regions), allocated via Territorial Just Transition Plans; its budget is EUR 17.5 billion for 2021–2027 (EUR 7.5bn under the MFF plus EUR 10bn under NextGenerationEU). The EITI earmarks transparency obligations at the extraction point; Superfund remediation trust sub-accounts are a parallel.
Co-occurring feature to add. The JTF pairs the earmarked sub-fund with a conditionality + plan-approval gate: countries access funds only by submitting an approved Territorial Just Transition Plan, and allocation is weighted toward the hardest-hit — per the EUR-Lex summary of Reg (EU) 2021/1056, "Only 50% of the national allocation will be available to those Member States that have not yet committed to" climate neutrality by 2050. Your sub-fund could condition disbursement on an approved transition plan for the targeted tier, with allocation weighted to vulnerability metrics.
Concerns with principle-level viability
The highest tax burden should fall on the largest market actors, likely with the best-qualified legal advisory and representative personnel at their disposal, consequently with the highest likelihood of unrealizable liabilities. So in its present form, this carries the risk of targeting precisely the least externality-causing segment of the economy (and as mentioned above, creating an administrative and infrastructural overhead to enforce that).
What if instead only the remediation fund is set up and contribution is voluntary? In the meantime, the overt non-contribution is an explicit signal of negligent operations.
The entities to whom this may give a real incentive to moot upgrading against the incurred externality cost are exactly the ones that, if they can be convinced that keeping up with the state of the art is vital, that will crowd out their judiciousness.
Could one incite the same FOMO on a low-externality life?
The reason why I think not is that SOTA following takes less attention and effort from executives than externality research, which tips the balance in favor of the former.[3]
Could the extra effort be outsourced, for a minuscule price perhaps, to a growth tax governance advisor body? Thus turning the decision into the trade-off between incurring externality debt and investing in better production?
What may the actual price tag be? This will end up pricing in the effectiveness of FOMO-inciting propaganda, not much different from e. g. AI labs' current attitude towards supporting AIS research and engineering.
On why I am admitting defeat in trying to articulate this
I want to first comment on a potential object-level debate of growth tax, then revisit key points.
The intent
Directly: to counter the investor optimism, funding influx and key role in stock indexes of IT in general and AI in particular, which enables the continued exponential growth in these areas.
Indirectly: (and objectively) to counter the opportunity, provided by rapid innovation and frequent but minor breakthrough in a technology, to support a public image or market position that the corresponding industry activity
is part of an ongoing revolution,
as such is expected to bring about key solutions to problems including societal ones,
lack of support e. g. from nation economy leads to falling behind and loss of sovereignty and security,
the activity, as both a driver of production and academic research, is central to preserving momentum of improving prosperity,
it is necessary for upholding economic traditions that have brought us flourishing at large --
-- whereas in fact these claims may all be ungrounded and create a false sense of value; at least comparatively to the externalities incurred by production.
Another important angle to the sort of progress (appearance) obsession that I have been trying to approach throughout this post is that incompleteness and in particular, vulnerability of recent and not SOTA products can be increasingly normalized or exploited to redistribute access to arbitrarily selected persons. For a maybe clearer and more detailed take read Notes from a Cannon Man.
Intentional simplicity?
I am admittedly not at the top of my public writing game, yet I felt urged to share this before waiting to improve. One shortcoming that I sense in particular is that I have been failing to balance simple, understandable reasoning with accuracy. I wish to convey a reading of my dilemma that it is actually quite simple but answering it seems to mandate ever more and more detailed qualification and breakdown, up to the point of realistically either only making marginal progress or not even moving in the direction of the simpler big picture. But leaving it unanswered just looks collectively suicidal.
Revisiting earlier highlights
In particular, dismissing the questions of Conclusion, or deferring to academic social scientific research or something similarly far from affecting population behavior directly, is dismissing that humanity may take lessons from history and foresee and avoid impending societal collapse. I believe furthermore that a diverse portfolio of talent should take them on and it is not a matter of qualifications but organizing the responding effort if a reassuring answer can be found.
Finally, now may be the best time to add to The presentation dilemma that with my present view, not trying to use my communication to persuade is a waste of credit and energy. I have spent my life deeming the persistent demand for objectivity and skepticism that characterizes science superior to other modes of thinking -- and by extension, rhetoric. Moreover, if I have to spell it out, I consider the present problem a scientific one. So ultimately this mix of perspectives is quite disorienting: the very drive to attack a challenge is exposed as contrary to a necessary ingredient of success.
as in: such an inherent feature or dogma of the culture in which a field lives that it robustly survives multiple paradigm shifts ↩︎
Both of these can of course be practically fulfilled by different instruments than taxes and here the usage of this term is primarily to symbolize this bundling of functions, rather than satisfying the related classical formal requirements. ↩︎
Plus perhaps a less likely contributor: being part of a crowd following the same trend gives (over)confidence in the safety of the individual (cf. the proverbial lemmings). ↩︎
The point. We all are confidently overlooking the root cause of why the alignment problem exists and I claim to have found a cluster of reasons but I cannot convince the right (amount) of people to take them seriously to make any progress on it. It relates to ML-involving technology at many levels of specificity and many points of the supply chain that enables it. (Also, please suspend your expectation that to meaningfully address a fundamental problem as this, using highly technical language or invoking high-context concepts is absolutely necessary, in fact I shall address this in particular On why I am admitting defeat in trying to articulate this.)
The short, poetic-intended version
With that, recall how the paperclip maximizer story used to be a go-to metaphor of how devastatingly risky it could be to unleash an ASI (if super-humanity ever was crucial to it in the first place) for 3 reasons at minimum:
and live happily ever after.Squinting at the paperclip maximizer, what does it let you see? As for me:
And of course, other details as well, which may or may not be necessary for that outcome, but the outcome is a universe consisting entirely of paperclips -- something meant to be absurdly ironic.
For what can this be a fitting metaphor?
Observe that it implies the belief in -- perhaps also a desire for -- a collective self-imposed ontological shift: the "we", who succeeds in engineering the form of intelligence that we are looking for here "correctly", is expecting that that very act is so rewarding as to never want to revert it. On the contrary, such success is an achievement that is instrumental to the prosperity of the collective, long term, so why would we?
An aside on mythologies
There are 2 assumptions that this takes vitally seriously, rooted in elements of culture, one more contemporary and another more ancient (perhaps archaic to some readers), respectively.
Firstly, the tenet that additional pieces of knowledge, behavioral traits, mechanisms etc. when viewed in a historical context consistently tend to improve the prosperity of the population. This is not really logically defendable if we adopt the position that ontologies shift throughout history and become increasingly incomparable. It does completely make sense as a belief that implies that the current living experience is not just what we can best relate to, but also the one we would be approximately the most grateful to have among all possibilities from history. It is a valuable mechanism to feel good about being grounded in physical reality.
The other is an important motif in the Judeo-Christian creation myth: any step of the creation starts with an act of separation that is contingent and not inherently permanent or even stable, and it ends with the meta-level observation that this is good. Accepting this creation myth not even at the level of making it part of a belief system, but just as a semantically correct and coherent text, asks for the context that the path from seeing that this is good to this being a permanent, even foundational feature of the universe is trivial. More explicitly: good should be what looks evidently, conspicuously in front of us and also that certitude and it being the basis of goodness should persist.
Observe furthermore that an incentive can be made polestar-like trivially as long as it involves an ideal that can be specified in ever-increasing precision and the deviation from that can be measured ever-increasingly precisely. I claim this trying not to restrict the definition of "precision in specification and measurement" arbitrarily at all. The best I can put it is: anything counts as higher precision if characterizing two things by the same scheme it (the more precise thing) overwhelmingly tends to take more information than the other, whatever the scheme.
Then note also that it seems intuitive that an incentive, like the incentives we humans like to follow, be best described just by some metric rather than criteria to meet. If we state something like "we just want to live healthier" but we can not quite delineate what that looks like, we can still likely come up with something to measure and by which to score our health, so obviously healthier means achieving a higher score. Wow -- this immediately even supplies a formulation of our goal that could be operationalized -- all we need is to play the game of pushing the score up and we should expect to live healthier! I think I am conveying here no more than the Bellman Q theorem.
Putting all of this together: what ticks boxes 0, 1 and 2 above is the very coordination that we see ourselves enact in the pursuit of ever-improving prosperity. I think we quite universally believe that
Conclusion
Something may seem conspicuously imprecise with the above analogy: I have not yet mentioned any way in which the paperclip itself is an appropriate element. I shall do that now by answering "What is the product, the creation of which we are assigning to superhuman coordination?"
I find the paperclip to be the symbol for something that is both directly conducive to reward and the end product of instrumental action sequences. If the subject, in our case, superhuman coordination has begun to drive the "health metric", which is designed to accelerate all of (a) some object-level production, (b) instrument design and (c) meta-level strategy improvement (i. e. learning), then an indicator that this has happened is a signal that is progressively increasing and its increase is accelerating roughly proportionally to its temporary value, i. e. and exponential signal. More simply I want to say that the subject having recognized and acting out an abstract, widely applicable self-improvement policy, it should robustly produce some exponential signal, and that is what the paperclip stands for. Conversely, the symbol could also be the moles of matter consumed in (paperclip) production (which in case paperclips are standard size is just proportional to paperclips produced).
A natural next question should be "Which signs can all this point to that expose that some form of superhuman coordination is on track both to preserve itself and to dangerously maintain the consumption of something essential to (human) life?" And personally I want to follow that up with some that were also instructive to ask when I first encountered the paperclip story
Progress status
If you're reading this, you likely know that the text is left as a draft, to put it bluntly, because for all my thinking and research my answer to "If we know what it takes to slow down a current paperclip-type superhuman coordination, then can we do that?" is a convinced no. And that made me pretty devastated and depressed and the reason I still decided to share this with you is that I've some hope left that you care and want to change my mind.
The (slightly) more detailed and quantified version
Take statements that roughly have the shape "Heavily based on empirics, and perhaps impossible to ever cleanly deduce from theoretical first principles, yet it seems to robustly hold that sociometrical or economical signal is exponential (in , i. e. time)." Or "Welfare is characterized by a robust >1 proportional growth over a sliding window, yielding the expectation of a corresponding signal." My favorite examples are
Now I want to look at some quantity like semiconductor market size over GDP per capita, to express the share of semiconductor industrial production in each person's contribution to the economy. More colloquially, I'd like to know the expected proportion of someone's added value coming from processing natural resources into chip substrates, and more concisely, estimate how much citizens are dedicated to wafer manufacturing on average. (Or, somewhat bluntly, our relative apparent need for silicon.) With all this elaboration I'm trying to be precise about my subject because I'm not sure if the quantity I end up investigating truly is the fitting and transparent choice, but more importantly, if I can convincingly argue that it is, both of which are crucial to make my point.
The presentation dilemma
The choice of quantity must be unbiased, to no extent cherry-picked to support my point, but honestly reflecting what I'm concerned with (again, quantifying our collective obsession with wafer production); and then it has to be meticulously researched, with the same intent. With that I haven't gone through, and it's worth mentioning why: as far as I can see, disentangling concurrent effect mechanisms that influence macroeconomic data series (especially ones that I have legal and / or physical access to) is rather complicated. Definitely so relative to the simplicity of the mathematics in the argument ("physicist's first approximation") that I wish to make. In addition, the argument is politically not neutral and nor is my personal affective attitude. Hence my motivation to see if I can make hard-science-supported, robust claims to argue for actionable, impactful interventions or shifts in culture. All that fits the type of manipulative propaganda all too well for my conscience.
To make matters worse I even strongly hold on to an intuition of what the shape of the simple mathematical result to derive and present should be.
The observation
I'm expecting the polymorphism, the wide-spread techno-optimism (or techno-fatalism?) of Moore's law and eagerness to extend its scope to various branches of IT and its supplying industry to manifest in several economic metrics growing in exponential order over time. Including ones that describe raw material extraction in particular and externalities or risks thereof in general. Meanwhile, GDP per capita (irrespective of subjective correlation to quality of life) stipulated to grow in exponential order and that being widely attempted to attain leads me to trust that such growth shall be observable long-term.
A proposal draft
"Taxation" in the name is not meant entirely literally: it has more to do with 2 notions referred to by the word [2]
I want to be upfront that I find it easier to conceive of my proposal as that of a concrete form of tax, however establishing the very machinery needed to enforce a tax is in tension (if not direct contradiction) with the intent of the proposed instrument. (I shall address this shortly, and most directly in The intent.) Although this section looks like a report of brainstorming the proposal, its relevance to this text is not its content but being an example of addressing a problem type. I did spend significant time and effort on coming up with details that I'll eagerly share and debate elsewhere if anyone signals interest.
Fundamentals
The effect mechanism: progressively with the adoption of upgraded, supposedly improved versions of hardware or software by participants of an economy they must contribute from their capital to the management of possible and eventual damages and externalities occurring at any point of the involved supply chain, e. g. by financing a remediation fund.
The operative domain: transactions in which a user first receives access to an upgraded piece of technology that they had a prior version of; the tax can be priced into the transaction value but it is negotiated on a per-customer basis.
Basis: the customer's activity, granting that their own production or service, social responsibility, criminal record, feedback for the developer etc. may warrant a reduced bar of entry (tax deduction) for the upgrade; however this must be arbitrated by an impartial body, e. g. an NGO that neither the producer, nor the vendor and nor the customer have common or conflicting interests with.
Particular addenda
Patches
This is a selection of ambiguous parts of the proposal that need precision and references to existing legal instruments that inspire in how to cover them (without claiming completeness).
Concerns with principle-level viability
On why I am admitting defeat in trying to articulate this
I want to first comment on a potential object-level debate of growth tax, then revisit key points.
The intent
Directly: to counter the investor optimism, funding influx and key role in stock indexes of IT in general and AI in particular, which enables the continued exponential growth in these areas.
Indirectly: (and objectively) to counter the opportunity, provided by rapid innovation and frequent but minor breakthrough in a technology, to support a public image or market position that the corresponding industry activity
Another important angle to the sort of progress (appearance) obsession that I have been trying to approach throughout this post is that incompleteness and in particular, vulnerability of recent and not SOTA products can be increasingly normalized or exploited to redistribute access to arbitrarily selected persons. For a maybe clearer and more detailed take read Notes from a Cannon Man.
Intentional simplicity?
I am admittedly not at the top of my public writing game, yet I felt urged to share this before waiting to improve. One shortcoming that I sense in particular is that I have been failing to balance simple, understandable reasoning with accuracy. I wish to convey a reading of my dilemma that it is actually quite simple but answering it seems to mandate ever more and more detailed qualification and breakdown, up to the point of realistically either only making marginal progress or not even moving in the direction of the simpler big picture. But leaving it unanswered just looks collectively suicidal.
Revisiting earlier highlights
In particular, dismissing the questions of Conclusion, or deferring to academic social scientific research or something similarly far from affecting population behavior directly, is dismissing that humanity may take lessons from history and foresee and avoid impending societal collapse. I believe furthermore that a diverse portfolio of talent should take them on and it is not a matter of qualifications but organizing the responding effort if a reassuring answer can be found.
Finally, now may be the best time to add to The presentation dilemma that with my present view, not trying to use my communication to persuade is a waste of credit and energy. I have spent my life deeming the persistent demand for objectivity and skepticism that characterizes science superior to other modes of thinking -- and by extension, rhetoric. Moreover, if I have to spell it out, I consider the present problem a scientific one. So ultimately this mix of perspectives is quite disorienting: the very drive to attack a challenge is exposed as contrary to a necessary ingredient of success.
as in: such an inherent feature or dogma of the culture in which a field lives that it robustly survives multiple paradigm shifts ↩︎
Both of these can of course be practically fulfilled by different instruments than taxes and here the usage of this term is primarily to symbolize this bundling of functions, rather than satisfying the related classical formal requirements. ↩︎
Plus perhaps a less likely contributor: being part of a crowd following the same trend gives (over)confidence in the safety of the individual (cf. the proverbial lemmings). ↩︎