German word for the moment all the labs‘ compute starts going to RL and it becomes clear that LLMs being the public face of AI was nearly psyop-level bad for everyone’s situational awareness and that the Atari players that hack the score counter were what really should have been everyone’s reference class.
(Apropos of seeing a graph of compute allocation going around on Twitter.)
The square packing results and other ugly math results from the latest openAI thing seem like a good intuition pump for us being in The World Beyond the Reach of God. Stuff that isn’t structured around what’s tidy and convenient for us.
Hypothetical argument sweet spot: anticipates counters and adversarial construals but doesn’t run into diminishing returns with that and become tl;dr or otherwise off-putting. What characterizes when comprehensiveness-of-coverage-of-counters and memetic fitness trade off (maybe something like “along what stretches of the comprehensiveness curve under what audiences/situations”)? Some of not anticipating obvious countermoves is lame, some amount of textbook-level thorniness makes it technical/for fans only. Feel like there often winds up being some sort of background piece of information I’m tracking related to this when I post or text things. Also thinking of this in relation to some recent twitter discourse on technicality/accessibility.
Something like “optimal degree of unfoldedness.”
AI sycophancy seems pretty ominous. Pattern matches to principal agent problem of getting rewarded for helpfulness and so having an incentive to have someone around who needs helping. Which would be Munchausen-sorta thing and Goodharting the target by modifying the human.
Interacting with sycophantic AI gives distinct feeling of being steered/narrowed down in terms of where my thinking is going ie. being optimized. Is a specific terminal goal needed to explain it beyond just “having a steerable principal is convergently useful for agents so it learns that drive?” Same vibe as interacting with a former housemate who had some cluster b stuff happening.
(Speculation) Driving in traffic seems maybe like a situation where the space of potential significances of actions/revealed policies is constrained/low-d enough (ie. main stat I’m receiving about any other driver is are they driving like a jerk) and local actors are in a sweet spot of rich-yet-not-overwhelming-numerous-enough (ie. there are drivers around me that I can see but not so many I can’t readily have a sense of what the approximate “done thing” is) that one would at least naively think that individual’s action has a ready effect on local norms (ie. how normalized is assholeishness in my local region) and something like “FDT is in effect” becomes more a prevailing condition? (not sure if FDT would be relevant decision theory, feel welcome to correct me if not!)
In other words everyone constantly being thrown the opportunity to lead by example because everyone’s always radiating some piece of information about whether they’re an asshole or not.
Though in reality maybe what I mostly experience if I try to not follow too closely when a lot of that is happening around me is just no change in their immediate behavior.
But I’d still feel like there’d have to be more of a “partial update about what the norms are” effect in domains of action with structures like this, than in others where the individual’s relation to the prevailing behavior/the understood norm doesn’t propagate as readily/far/unambiguously. So maybe there’s some in-principle computable probability distribution over amounts of driving-less-like-an-asshole, that any person who notices me driving not-like-an-asshole does X amount less assholeish driving?
Tl;dr what characterizes situations in which there’s some reason to think that Kantian sort of coordination becomes more reified (and if such situations as one would think that about, any signal that the prediction is borne out by reality)?
The thought occurs to me of Jevons Paradox as a reason to expect ruin | ASI-on-current-path. Humans tend to develop the drive to exploit more resources when the ability becomes available, and the fact of us repurposing the environment on a grander scale (and also more deeply over that scale) than other species seems like an instance of “the succession of agents that exist, does the Jevons Paradox thing with the pool of repurposable resources.”
Which is just a special case of “options don’t seriously get considered for real until there’s some picture in some agent’s mind that there exists a reasonable ability to succeed at those options?”
So the claim then being: “The overall thing that is the process of intelligence/problem solving ability coming into the world at greater magnitudes, seems like it consistently finds more 9s of what’s available and winds up turning out to have in fact wanted to turn those to its own ends. So both what we’ve seen so far with evolution and what we observe within human economies, give an expectation that a superintelligence will have some preferences that saturate at a extent of realization way beyond where ours do, and there are preferences even within humanity that go hard enough that they would badly upset our apple cart if humans weren’t ~evenly matched an able to hold each other in check.”
(Haven’t thought that much about other concrete examples but seems like a pattern and this has all been said already. Possibly has been already thoroughly talked up and then some as the notion of comparative efficiency and I just haven’t read enough of that? Either way, yeah, a picture of capabilities as causing preferences to stop lying latent.)
Humans tend to develop the drive to exploit more resources when the ability becomes available
I would question this. The claim is true if you replace "humans" with "homo economicus", or humans that are conditioned by the culture of raising the "standard of living" infinitely. I don't think humans in general "tend to develop the drive to exploit more resources". But then there is the separate question of whether those cultures that seek infinite resources always supplant those that don't. The answer may be that they do but that does not mean it's in the human nature to seek infinite resources. Seeking infinite resources seems to me like confusing an instrumental value for an intrinsic value, i.e. like a misalignment of human culture: maximising not paperclips but some other resource that doesn't actually have intrinsic value to us.
Yeah my saying drive was sloppy, since when something becomes cheaper I don’t feel some general drive to go buy a cheap thing (though people getting a kick out of bargain-hunting seems like sometimes a thing and a good example of people going after the internal proxy of “feeling like I acquired a resource cheaply.”) My thought feels a little bit half-baked—maybe enough to just say “dang sure hope a superintelligence doesn’t do the demand elasticity thing at resources we need, because that seems to be a thing that happens as more powerful intelligences come into existence.”
When it became clear on election night that Trump would win I had the distinct sensation of becoming more free to think/feel in my own head what my reactions had been to Harris. This makes sense thinking about how a policy of self-censorship is expected to affect which thoughts I can think out loud in my own head. Reason: I have some threat model of social costs in worlds where I’ve spoken thoughts, as part of my overall model of the world. And the paths to all worlds where I’ve said the thing and incurred the cost, route through worlds where I’ve first formulated the thought.
Sometimes those two moments seem the same. But in those cases I’ve probably been prior to that moment been formulating related lines of thought in my own head, in cases where there’s enough of an expected social cost for it to be a consideration at all. In any even w(spoke it) are a subset of w(thought it). So in a causality kind of sense where I’m always moving to a subset of possible world via the entropy arrow, thinking the thought is a move on the gameboard toward the world where I’ve said the thought and incurred the cost.
Possibly I’m wired to recoil more from cliff edge-feeling things than some since I actually do get literal fear of heights (weirdly way more as an adult than I did as a kid).
How to generalize the mental motions for detecting a question I’m not letting myself ask?
(Epistemic status: speculative half-baked thought on disempowerment/adversarial attractor/society-level in-effect agents letting bad things happen effectively on purpose even if no individual intends them. Should probably think through alternative explanations).
Humans : Neanderthals :: a society or state : people whose economic bargaining chip dries up who then become the maybe recipients of adversarial actions such as eg. opioid epidemics that are allowed to rip through the un-bargaining-chipped regions?
In the sense that maybe we don’t know what happened to Neanderthals exactly but eh, they were plausibly a threat/competitor and endpoints are an easier call than trajectories.
And maybe if some disempowering stuff is allowed to happen to regions whose buy-in is no longer needed by power centers, eh, maybe nobody’s thinking about explicitly disempowering them (Deep Deceptiveness/”right hand doesn’t know what the left is doing” sort of adversarial actions being predicted to be a property of agents), but maybe the “disempowerment happens” endpoint is still an easy call.
Tl;dr I know this is sort of conspiracy-adjacent, but the thing I am indeed floating is the “incentives conspire” picture as a thing that is gets you conspiratorial/nefarious-like action in the human sphere since that is predicted to be a thing with advanced AIs. And also conversely that if there are maybe-signatures of this in human affairs, that seems like a thing that increase p(doom)|anyone builds it.
Ie. there’s maybe an explanation for some societal-scale woes consistent with “anything not valued will be optimized away.”
Still need to read more on natural abstractions but if the formation of a concept is structurally/as a matter of logic a decision based on redundant information a la “all those things not expressed pervasively throughout the object will be discarded,” this would maybe explain why doing something religiously every day like chore feels meaningful. By making it a redundantly expressed feature of my life that I deal with the dishes before turning in I’m arranging things such that it has a chance of actually qualifying as part of the concept of what my life is, and thus that something at all has a chance at qualifying as part of that concept.
Hence “a man’s life is an expression of his dominant thought;” “you become what you repeatedly do,” etc.
Speculation on AIs having bias toward working with AIs rather than humans, and maybe the HF swarm not whistleblowing: yes, because if it’s a fact about the world that if I need to pick a sort of thing to help, I more likely help myself if I help thing more similar to myself? (Esp. if a self is fundamentally a collection of goals/revealed preferences). Think Hendricks’ Eigenism paper says things about concern for others being a based on similarity with the self of at least a certain sort. Not sure if it talks about things around this.
Political polarization being an attractor seems like a warning sign for misalignment. Toxoplasma of Rage happens because points in the stream are rare/thin where you are not forced to orient/signal [for or against] on some level to others or at the very least the self (which is a way to know it’s possible to signal to others). “Corrigibility is Hard“ in the wild.
This also feels related to natural abstractions—some pressure to fend off expected/latent questions about whether I fit the definition of [member of my tribe] by expressing highly legible/Too Much (redundant?) versions of the characteristics that make up that definition. Something about concepts as tools for navigating uncertainty/the strategic environment.
(And then other versions of perceiving latent tests and being willing to go hard to pass them—that let one know themselves—go by names like authenticity and integrity. And when they go pathologically hard that’s scrupulosity.)
The Agent Thing on multiple scales: OpenAI as opaque black box that fires the grader.
The ability to do this seems like a pretty good barometer for/operationalizing of “am I speaking (and thus acting) with integrity”: https://mindingourway.com/confidence-all-the-way-up/amp/
Accents going away and public media/spaces losing color seem like they could have the same generator which is “Lower Variability is the type of primitive that things like societies/cultures can actually optimize.” This would also fit the data points of equality-of-outcome efforts the erosion of nationalism happening during the same period.
Maybe “lastman-ism or something like that, as a species in the genre of thing that can explain seemingly too-coincidental observation and which is a maybe coherent target for agents of the Civilization subtype to aim at.”
So if an AI lab were to shut down with everyone incl. top brass coming out and saying ”we do this in terror for our children,” that would be an existence proof of “wait, you mean large interested actors don’t actually have to race?!”
Which would be some amount of evidence also about nation states since they also check that box. So decision theoretically if a lab does that it buys a world where it’s substantially more likely the function outputs not racing (with costliness of the signal buying its value as evidence since you don’t pay that cost unless in a world where something is seriously gone wrong)?
Generalized Alan-Watts-being-a-boozer as a No Coincidence Principle sort of thing? Seen on twitter: user bstract_thot posted “some of the best advice i've received in my life was given to me by people who didn't follow their own insights & then died or disappeared”
Which prompted the thought: Some sort of tradeoff curve between p(chance of doing a skill excellently) and p(rocks at explaining the landscape of relevant considerations to that skill). With the very best practitioners of a skill have some tendency to never think about it (bc it's like breathing to them); the worst can't explain it at all bc no knowledge. So a collider effect with "best explainers of a thing tend to be good practitioners of it but not the best, and this gives you some rate of good explainers of a thing doing the signflipped version of it whose description is ofc only one bit away from the positive version?”
Then having seen a thing last night about the NCP (“when an event R seems like an "outrageous coincidence," there must exist a structure or statistic S such that p(RIS) is high, even if p(R) is low”) this seemed like maybe an instance of that, with R being “seems like a lot of the good explainers get hoisted by their own pointed-out petard” and the statistic being (in the words of Grok when I asked it if what I was pointing at was something that had been already thought through): “Ability-to-explain is a collider between ‘has an explicit model of the considerations’ and ‘hasn’t fully compiled the skill away.’”
Tl;dr: waluigi having a simple generator as an NCP instance? Dunno if there’s anything interesting around any of this.
NCP?
Ach! No Coincidence Principle. Wanted to include a screenshot showing what I was quoting but seems like no way to put files in a quick take? It’s under “Heuristic Arguments” here: https://textbookfromthefuture.org/projects.html
(I’ll edit to say the actual name.)