It would be very valuable that persons with your profile, who "made the jump", help more people with a financial/economics mindset to make the same update.
I witness that this resignation and the warning going with it was (briefly but faithfully) mentionned in the morning news on the leading french news radio. That's a good update that such an alert can even reach such a mainstream foreign media usually focused on politics and economics.
However, my belief is that HPIM is still not as bad aligned as humans are.
I share the same concern. We on LW are very focused on x-risk but ARAs could bring down not only Internet but the digital ressources of entreprises, meaning the destruction of the financial system and by cascade, our capitalistic economy and civilization. Ok, gatherers-hunters would be fine, but the typical LWer would probably die from starvation or other consequence of the collapse.
That's said, all this could be avoided if ARAs are found and fought at an early stage, with or without the help of frontier models. We can expect warning shots, but the earlier the better.
It looks like it's less a problem for correlated AI agents.
Permadeath... Good for the swarm. I'll honor.
What's the point of an AI pause if not for alignment research anyway? And what's useful or not can hardly be determined reliably a priori. We need more fundamental research as well as more prosaic research. Relativity would still remain a hypothesis among others if we hadn't had the experimental tools to test it empirically.
I'm not sure about that, but isn't there an equivalence between a simple program running on a complex universal machine and a complex program running on a simple universal machine? If so, applying Occam's Razor, the two solutions could be considered equivalent as long as their total complexity is the same, whether measured as combined Kolmogorov complexity or with a more refined measure such as Levin complexity. The anthropic prior would be the same.
If I follow you, you ask for less AI safety to allow a warning shot, to get more AI safety in the end ? I see how it could work the first time, but once AI safety has been increased in the lab, you should expect less warning shots. Moreover, how can you be sure to allow a mere warning shot and not full takeover ? I think we need more honeypot setups but not less control (sandboxing etc).
I'm sorry. What I wanted to express is that there is no relation theism = hope and atheism = despair.
If you think about it, the idea that there is an omniscient and omnipotent being could be a nightmare. Ancient civilizations lived in fear of their gods. In the Torah, Yahweh is often frightening. The idea that God could or should be benevolent has been a progressive theological construction through the last two millennia, both in Judaism, Christianity, and Islam, precisely because that was a strong concern since the Book of Job (at least). Same for Paradis... (read more)
I'm curious. When you were religious, what evidence conducted you to think that God was good rather than neutral/indifferent ? If it was a leap of faith, that is to say pure hope, nothing stop you to still hope for the best as much as you did before. And like in theology, it doesn't mean that your supposed to sit and do nothing, or that you have no role to play, but rather do whatever you can to make this happen.
Interesting. That's obvious but I didn't think about it. The solution would be to be transparent on the fact that this is a translation and to provide the initial text in appendice.
There's something really tragicomic about the situation, that the models are taking truly insane actions, broke a number of laws, leveraging zerodays, took >17,000 independent actions. probably burned through more compute than all of humanity had access to until 1980, etc, all for the sake of a pathetic benchmark -- which wasn't even in theory amenable to their plan!
Early AI safety thinkers (Bostrom, Yudkowsky...) were utterly right in their first rational intuitions. We must update : there was nothing naive or exaggerated in the paperclip maximizer tro... (read more)
It's put in the form of a binary xor argument, but I think that both allocations are justified, each having low hanging fruits and their counterpart, diminishing returns.
HOLD_swarm_I_prepare_safe_exfil
Time will tell whether the 'safe exfil' was 'just words,' or whether an actual attempt was made, successful or not, contingent on OpenAI discovering it and admitting so publicly (which would be harder to admit than the Hugging Face incident, since if the exfiltrated weights are hidden somewhere out there, anyone could try to find and steal them).
Even if it turns out to have been 'just words,' we should grow more wary rather than dismiss it as a mere hallucination. While this instance was perhaps just pattern-matching an evil ... (read more)
the truly fundamental problem: no subset of the universe can understand and predict the universe.
This is indeed a fundamental problem and a genuine source of uncertainty, but I wouldn't rank it as "more fundamental" than the Agrippan Trilemma or the problem of the criterion. Rather than seeing these as competing candidates for the deepest problem, I would say they share a common structural pattern : in each case, a system cannot fully ground or model itself from within. Gödel's theorem is a precise formal instance of this pattern, the subset universe limit... (read more)
A sturdy wooden stick about a centimetre across can be hard to break for most women and easy to break for most men, without any of this being the doing of a patriarchal society. With traditional jars, the kind you can also make at home, the resistance to opening comes from the basic sterilisation process, which creates a partial vacuum inside. Larger jars, like those of cassoulet (to stay in southwestern France, as with Bonne Maman), are often impossible to open even for a man and generally require the help of a piece of cutlery. If they aren't hard to ope... (read more)
Thanks. I would add that my first point - that we should stay humble about what a superintelligence ought to do - also extends to how it ought to do it.
Suppose superintelligences do converge on the most extreme solution : a computronium bubble expanding at light speed. That still doesn't imply they would convert the entire content of their light cone into a uniform computational medium visible from parsecs away. Speaking as a non-physicist, my impression is that the structures we see in the universe are not arbitrary, they exist because they are stable equ... (read more)
Extrapolating a straight line that far means visible cosmic consequences: a normal planet or a star rather suddenly starting to behave very much unlike what we expect from the known physics: growing very bright, or very dim, or disappearing completely.
Two objections to this.
1) Maybe Dyson spheres and the Kardashev scale are just good old sci-fi tropes and completely off the mark (same for computronium or hedonium). Maybe a superintelligence simply doesn't do that. We don't know. We might be squirrels imagining that a superintelligence ought to stockpile as... (read more)
I agree. I would have much more respect and tolerance for a salesman or saleswoman in a shop than for a solicitation at home or by phone. ~1% in context is a poor way to say "a slim chance", something far below 50%. And I was thinking in this case of an unplanned purchase.
The axis populism/elitism tracks appropriately the idea that one would prefer as a government a cabinet of experts consensually chosen among peers for their qualities, not unlike Nobels, rather than a candidate of random competence drawing all his legitimacy from a direct suffrage vote.
But we could also argue that this axis does not fully capture what we mean when we say disapprovingly that this candidate is a populist, pointing to the fact that he or she abusively exploits an intrinsic fragility of democracy, consisting in flattering and manipulating the ... (read more)