Good post, but I am curious about your thinking for :
> when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
Isn't this the small world case? Following the pattern, my big world example would be: "when there are so many people around that you rarely meet the same person twice, you should scam people as much as you can". Or do I have a different interpretation of small/big world?
Interesting, your small/big-world point reminds me of the mean-field footnote in OP that I felt confused by:
There are zones where the best move is non-consequentialist, in the sense that you do best by following heuristics (or principles, or virtues) that are good regardless of the exact specifics of the situation you’re in. You can neglect your (immediate) action’s (immediate and specific) effect on the environment.[1] [1]: This intro to mean field theory seems to be pointing to a similar concept of a “foreground” particle and a “background”, where the background is made of the interaction of lots of little “foreground” particles, but each particle has negligible impact on the background; you can just model the particle’s behavior as some function of the (assumed approximately fixed) background. But I don’t really know mean field theory in physics so I can’t be sure the analogy is apt.
Here, assuming that I have no effect on the environment is not quite what mean field theory does (at least not in the places where I encountered it in physics): mean field theory assumes "everyone behaves like me. If I follow a principle, the world does too." From that assumption, a natural next step is Kant's 'Act only according to that maxim whereby you can at the same time will that it should become a universal law.'. This does seem related to "the best move is non-consequentialist".
If I had zero effect on the big world (also in the sense that the big world does not treat me any differently if I misbehave) there would be no force that pushes me to cooperate. But with the mean-field worldview, I will assume that there is an effect even if I am not tracking the causality – whether there are repeated interactions or not.
the "world" i was thinking of was the iterated "game" of interacting with people. This world is "big" in time because you expect it to go on for a long time in future. So the immediate impact of your next move in the game is tiny in comparison to long run effects like your reputation or your habits.
And how you behave determines the size of the world you can play in! Are you part of the vast virtuous mutually normative collective of moral agents that is our society? Or, as often happens with smart self-interested sociopaths, does your game end early in jail or shunned or dead.
I hesitate to curate this, but am curating it.
I hesitate because this post raises the question "our usual heuristics on how to be a good/effective person don't hold up during extreme circumstances. wat do?". The post then doesn't answer that question.
I'm scared of many people reading this, thinking "I'm an exceptional person during what is maybe the AI Midgame, maybe approaching The Endgame, and indeed yeah the commonsense morality/heuristics don't make sense during these extreme circumstances." I am currently way more worried abut those people jettisoning their morality without thinking it through, than them failing to handle nuanced Midgame/Endgame situations optimally.
The reason I am curating this post is because I think it does a good job of laying out why you should have a lot of commonsense intuitions on how to coordinate and participate in the world, in a way that lays out a different set of mechanistic gears than I've seen laid out before. This is both conceptually interesting, and maybe motivating to some readers who hadn't realized why this was important.
I think we have at least awhile more than all the Big World heuristics just straightforwardly apply, even if the circumstances are starting to feel extreme. In the end, my guess is when you are sufficiently galaxy brained, it turns out that basically you want to just keep applying these intuitions even in various extreme circumstances. But I don't have a succinct argument for that, and would like to have more clear and robust arguments for that, uh, soon.
I'm scared of many people reading this, thinking "I'm an exceptional person during what is maybe the AI Midgame, maybe approaching The Endgame, and indeed yeah the commonsense morality/heuristics don't make sense during these extreme circumstances." I am currently way more worried abut those people jettisoning their morality without thinking it through, than them failing to handle nuanced Midgame/Endgame situations optimally.
This part of the thought process itself feels not very Big-World-pilled to me? Like, in a big world it's generally a good thing to promote interesting/valuable ideas, even if a few people might misinterpret them and go on to do bad things on that basis, or something?
I'm not even sure these heuristics break down mid- Endgame. Like, when does gathering resources become stale? If you have 5 days left to live, you need water still.
And note Sarah's "big world" example of meeting n=small people, n=large many times. There was still a Big World heuristic to extrapolate. So that's interesting, it means I like this post, but I also notice that "modeling a big world" is left as exercise.
I also just directionally agree that prematurely applying Endgame thinking is bad. This keeps coming up on LessWrong.
I think the important question is more like, when you think you're Small world, find or model the Big World. It's not "oh we really are all dead this time," it's, "there are probably some reasons to keep behaving sanely that are not instantly available to me right now and I should find them."
when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
I'm not sure why exactly you believe this, can you say more? (Calling the arguments for backfire "galaxy-brained" kinda unduly stacks the deck, I think.)
Heuristics like "doing a good job on important work benefits the world" and "it's beneficial to share scientific knowledge freely" are time-tested (at least over the past several centuries) and have a lot of empirical evidence in their favor. "The Baconian project has been working" is one of the only sweeping historical/social claims I am highly confident in. And you don't even need that to believe a more common-sense thing like "working to solve problems that people find helpful to have solved is beneficial on net." Like, "productivity is good" is something that could have been intuitive in antiquity or prehistory. There's just tons of individual cases, life experience, and logical extrapolation that all point in the same direction.
"THIS time if I invent/improve this technology it'll be harmful" or "THIS time if I disseminate this discovery it'll be harmful" fly in the face of those heuristics and depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it's a fairly conjunctive argument full of lots of beliefs about what people will do, made without specific knowledge of the kinds of people in question. Doesn't mean it's wrong, but it's risky.
Thanks for clarifying!
My worry is, these heuristics are only "time-tested" in that they've (apparently) had good consequences relative to not-super-large-scale goals, in not-super-unfamiliar contexts. Many people working on the kinds of problems you're pointing at do so for impartial altruistic reasons. So I suspect that that track record is only weak evidence that such heuristics will work relative to their goals, in unfamiliar contexts — like, say, ASI takeoff.
"Sharing scientific knowledge freely" contributed to the industrial revolution, for example, without which factory farming wouldn't have happened, nor the development of technologies that pose x-risks. Granted, the implications of these examples are debatable. But they seem prima facie concerning enough to me that calling big-world heuristics "robust to uncertainty" sounds way too strong (for impartial altruistic goals at least).
depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it's a fairly conjunctive argument full of lots of beliefs about what people will do
I don't think the extrapolation from "good on local scales / familiar contexts" to "good on the cosmic scale / very unfamiliar contexts" escapes really conjunctive causal mechanisms. It just hides them at a higher level of abstraction. When making that extrapolation, we're predicting that the same sorts of mechanisms that led to good consequences on the local scales will be at play on the cosmic scale. So we're baking in a lot of implicit beliefs about those mechanisms.
This reminds me of when Ada Palmer asked "What would Machiavelli say if you asked him what would happen if Milan suddenly changed from a monarchal duchy to a republic?":
The poli sci students went first: He’d say that it would be very unstable, because the people don’t have a republican tradition, so lots of ambitious families would be tempted to try to take over, so you’d have to get rid of those ambitious families, like the example Livy gives of killing the sons of Brutus in the Roman Republic, and you would have to work hard to get the people passionately invested in the new republican institutions, or they wouldn’t stand by them when the going gets tough or conquerors threaten. It was a great answer. Then my students replied: He’d say it would all depend on whether Cardinal Ascanio Visconti Sforza is or isn’t in the inner circle of the current pope, how badly the Orsini-Colonna feud is raging, whether politics in Florence is stable enough for the Medici to aid Milan’s defenses, and whether Emperor Maximilian is free to defend Milan or too busy dealing with Vladislaus of Hungary. “And I think I’d have something to say about it!” added my fearsome Caterina Sforza; “And me,” added my ominously smiling King Charles. In fact, my class had given a silent answer before anyone spoke, since the instant they heard the phrase, “if Milan became a republic,” all my students had turned as a body to stare at our King Charles with trepidation, with a couple of glances for our Ascanio Visconti Sforza. It was a completely different answer from the other class’s, but the thing that made the moment magical is that both were right.
when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
Isn't this one more of a "small-world" effect? In a big world, you're unlikely to run into the same people over and over and one-shot dynamics become more prevalent.
By contrast, if you are in the endgame of a game of chess, or in an oligopolistic competition, you absolutely need to think about how other players will respond to your actions. You do need to search through a tree of “if I do this, that will happen” and compute that specifically for the specific game state you’re in, and recompute it every time the game state changes.
Since chess is bounded you really can calculate (and in the endgame you're now able to). With oligopolistic competition, your power has increased and the players that matter to you are fewer, but the moves are not bounded -- if Apple and Dell (or whoever) are in an oligopoly but then Apple creates the iPhone...things can change.
Most of the time, for most people, this is not the case!
I mean it is interesting that we are, all of us, now in this case. And I do notice that it causes heuristic problems for even very thoughtful people.
This overlooks one important distinction: most of these scenarios focus on ignoring indirect effects when the direct effects are small enough not to upset the environment and change other parties' behavior. That makes sense, but the direct effects still need to be positive. Your ideal game of chess might not vary much with your opponent's moves, but it does not involve marching your king forwards as fast as possible. Likewise,
- when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
if you're doing AI research, it might not be significant how other people respond to your incremental progress, in an environment where many others are making incremental progress, but it is still significant whether making progress is the right thing to do. That is a separate discussion and not one that the pattern outlined here pertains to.
This intro to mean field theory seems to be pointing to a similar concept of a “foreground” particle and a “background”, where the background is made of the interaction of lots of little “foreground” particles, but each particle has negligible impact on the background; you can just model the particle’s behavior as some function of the (assumed approximately fixed) background. But I don’t really know mean field theory in physics so I can’t be sure the analogy is apt.
I think the more precise formalism here is mean-field games, which studies how a large population of strategic agents can be approximated by a representative agent interacting with a mean-field distribution of the population. This maps onto individuals determining their policy with a "big-world" intuition that their actions will not change the larger distribution which is generally at equilibrium.
In board games with two players reducing the opponents points is as good as increasing your own. In board games with 4 or more players the best strategy is to increase your own score even if it increases someone elses'. This is the many players strategy version of his big world heuristic.
I feel like you're saying that for two players chess because of how the evaluation bar looks like, but in a four players situation, you'd rather have four independent evaluation bars.
I.e you're making two different claims on two different methods of evaluation.
The statement "increasing your score is as good as decreasing your opponent's" is true by definition for the first method of evaluation in two players chess. And it has an equally banal generalization for multi-player board games (in which that evaluation method applies). That would be the correct multi-player version of the first statement.
The statement "increasing your score even if it increases your opponent's" is not true by definition for the second method of evaluation (for any number of players) but isn't even well-defined for the first method of evaluation.
So the second statement is not really a generalization of the first.
First, your small/large framework may need more conditions to "compactification“. Taking traders as an example, you don't really need to consider the impact of your actions on the market, but in order to buy at a cheap price, you do need to consider the impact of large players on the market, because good prices are basically the result of the actions of large players (most evident in the futures market). Therefore, strategic considerations have not disappeared; they have simply emerged from a different level and perspective. The "small/big" framework "compactification“ not only requires a sufficiently large environment, but also a sufficiently uniform distribution. Alternatively, replacing the small/big framework with a principle/strategy framework might be more relevant to your example. Principles are strategies in infinite games, while strategies are principles specific to a particular environment.Taking the example of" running into the same people" as an example," big in time" is essentially a problem of finite-time games (a single transaction) and infinite-time games.
Secondly, the principle that really needs to be considered for the AGI problem is the principle of "fault tolerance"—is the society bigger than the large language model enough to absorb possible errors?
Finally, the example I gave of the futures market may be a bit special, but the small/big framework holds true in many situations.
I think the chess example is inaccurate, but it doesn't refute the main message (or how I choose to interpret it at the very least).
It's very wrong to say that in chess, in the opening moves, both players do their own little thing and after some development, only then do the cordialities end.
In the opening phase, both players aim to achieve certain advantages and if you let your opponent get all the advantages they want, you won't have a very pleasant middle game.
Openings have a 'logic'. If white plays 1.e4 then black doesn't want white to get the full center and thus plays a move that prevents that, while also grabbing for center. This is the logic of 1.e5. Now white can't have the whole center and so focuses on developing their pieces. Maybe they could do it in a way that would let them keep the initiative and so plays 2.Nf3, getting the Knight out while attacking the e5 pawn and thus again forcing black to respond etc.
If you don't understand your opening's 'logic' and play against someone who does, it's very easy to get disadvantaged early on, possibly several moves before even noticing you're on the backfoot.
On the other hand, what is true, is the slow, gradual, incremental nature of 'advantage' in chess. Unless both opponents decide to go for a very sharp and complicated game (I'm assuming good players btw), then progress is going to be very slow and players won't really be thinking about checkmate until much much later.
This is the nature of Strategy vs Tactics. Tactics are about identifying a short, exact sequence of moves that leads to an immediate concrete advantage.
Strategy is essentially about developing a situation that makes it easier/likelier for there to be tactics for you to execute.
I think for people who habitually over-focus on details and worry about losing control (I guess many people are like this, myself included), the big-world mode is a beneficial and energy-saving lever.
Meanwhile, I think it is both the most inspiring and the most easily misused part of this framework. Something like “ Concentrate on your own work and remember you’re only a small part of the whole” can be both a healthy expression of the big-world mindset and the standard line of the thoughtlessness Arendt described . (To be clear, I’m not equating the two. The problem is precisely that it is difficult to distinguish between the two based on the wording alone.) Since the cost of misuse can be severe, the framework seems to lack a more solid boundary here, to keep its users from slipping from the former into the latter.
A mere “time-tested general rule” may not be enough, even without pushing to cosmic scales. Although we share most basic common sense, the “general principles” we each take for granted in belief and in practice can be different, especially when they involve decisions over others. And the track record of a “time-tested rule” is often written by those who maintain it, while the costs are borne by those who have little say in the matter.
Take stacked ranking as an example: for management , it simplified evaluation and “kept the company competitive” GE ran it for decades, it can truly be called “time-tested” while for employees, it fostered internal rivalry and punished collaboration. Before it’s clear what counts as a “time-tested general rule” and who did the testing, both sides have reason to believe they are the ones following big-world principles.
You mention at the end about the lack of thinking tools for endgame situations — admittedly, that’s an unprecedented and tricky problem for anyone. What I want to ask is a prior question: is there anything within the framework that could help someone recognize when it’s time to shift to small-world mode (or at least some warning signs)? In my view, that might be the more everyday, and more easily overlooked failure.
Any 'endgame' in the real world is an induction trap if it fails to acknowledge the reality that some situiations continue to infinity and do not have an 'end state'. treating such situations as an 'end state' introduces logical fallacies and assumptions casuing repeated downstream failures. While, I can not speak for others, I can say that I am endeavoring to remove these induction traps from my thinking as it is currently the path of least resistance to achieveing the goal of presenting myself as 'less wrong'. This is relevant context becasue my presentation endeavors to be an honest representation of what is on the 'inside'. its not a mask. its an attempt at a factual representation.
As for your thinking on 'when you expect to run into other people', I believe, you are correct and the tit for tat version of the prisoner's dilemma indicates the math favors trying. (It supports your statement).
I feel like "Big-World Intuitions" implies that the main breaking point of those intuitions is the size of the environment and this can be misleading. In one of the examples given: "share scientific knowledge freely" the intuition might break down even if you are in a Big-World situation. You might discover information hazards accidentally anyway. For those cases, it is better to share with specific institutions first before revealing to the wider public. Similarly, in the chess example: how you develop in the opening might depend on what your opponent is doing. If he tries to play scholar's mate you need to defend in a specific way by move 3, so even in the very beginning ignoring the other player is not an option. My point is that your relative capabilities and the scale of the environment are both relevant, but should not be the only factors determining when these intuitions are actually reliable.
I don't think it's the 'size of the world' that matters but the distance induced by how easy it is to reach and have an impact on other individuals. And so 'Big' is meant in the sense of this metric, not necessarily physical reality.
For instance, I don't think this heuristic applies very well to Go, despite the enormous size of the board, because players' stones are pretty much at distance zero. In fact, because I know close to nothing about Go, the above statement is a prediction based on my mental model. I'd be interested if a competent Go player could confirm or deny it.
I want to pattern-match this with Kegan 3 intuition (big world) vs. Kegan 4 systems (small world). Idea being: our brains are single human-scaled, with a fittingly-limited context window, which is highly compatible with broad heuristics and intuitions. As our power increases, though, and our actions can scale profoundly far beyond ourselves, the needed level of complexity and specificity is beyond the ability of our intuition. Make sense then to reach for more powerful computation: maybe some computer modeling to work out that "if this, then that, then that" chain. Lacking that, I think people in positions of great power get feedback another way e.g. public outrage, failing business objectives, legal battles.
Consider the following situations:
when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything adversarial “against” them, but just to focus on growing and providing value to your own customers.
when you are a small trader in a big market, you don’t need to worry about your trades shifting the market price or revealing information to your competitors; in many contexts, your optimal strategy is simply to bid your true price, buying when an asset is cheaper than your “happy price” and selling when it’s more expensive.
when you are in the early stages of a game, often your best strategy is to grow your “resources” (like developing your pieces in chess, trying to control more territory and have more value on the board), following a pattern that’s mostly independent of what the other players are doing and gets you more of something that’s valuable across many possible game states.
when you are a species whose resource needs are much smaller than the carrying capacity of your environment, you are r-selected; your fitness is maximized by just having as many offspring as possible and not “worrying” about running out of resources.
when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
These kinds of intuitions are what I’d call “big-world” heuristics.
What they have in common is that the world is “big” enough relative to you — the market is much bigger than you, the problem is much bigger than your progress on it, the game is very far from being over, the environment is much bigger than your resource consumption, etc — that your best strategy for achieving your goals is not to explicitly calculate how your action will affect the environment (including how your actions will affect other agents’ actions.) There are zones where the best move is non-consequentialist, in the sense that you do best by following heuristics (or principles, or virtues) that are good regardless of the exact specifics of the situation you’re in. You can neglect your (immediate) action’s (immediate and specific) effect on the environment.1
By contrast, if you are in the endgame of a game of chess, or in an oligopolistic competition, you absolutely need to think about how other players will respond to your actions. You do need to search through a tree of “if I do this, that will happen” and compute that specifically for the specific game state you’re in, and recompute it every time the game state changes.
Big-world intuitions feel a bit like “play fair, mind your own business, cultivate your own garden, do your best, and it’ll all more-or-less work out for the best in the long run”. Or like “keep your eyes on your own paper”, don’t try to manipulate people, just follow the same path you would if you were alone; whether you’re Robinson Crusoe alone on an island or one anonymous citizen in a big city, your “job” is pretty much the same, in that you need to work to take care of your own needs.
This came to mind because I noticed that I pretty much always rely on big-world intuitions. I don’t really know what to do about questions like “what if we win too much and it’s bad” or “what if I, personally, have the power to shape society, how would I choose to set it up”. And I don’t really even know where to begin with chains of adversarial strategic thinking like “if I do this, she’ll do that, so I’ll do this…” I’m always thinking of myself as one participant among many, with a small share of power/influence, working on problems hard enough that the only thing I have to worry about is doing the best job I can, and trying to follow universally sound principles/heuristics.
The nice thing about sticking to principles/heuristics is that it’s robust to uncertainty. What if you’ve misread the situation? What if you’re not as big a deal as you think you are? Behave in a way that’s usually for the best, even so.
Except, of course, if you really are super powerful, super close to winning, super close to the “endgame” in some sense, or in a really bizarre situation where the consequences of following generally-good heuristics happen to be very bad. Most of the time, for most people, this is not the case! But it is disquieting to notice that none of my grounding intuitions or usual ways of thinking about what “healthy” or “prudent” or “ethical” look like, are fit for such situations.
This intro to mean field theory seems to be pointing to a similar concept of a “foreground” particle and a “background”, where the background is made of the interaction of lots of little “foreground” particles, but each particle has negligible impact on the background; you can just model the particle’s behavior as some function of the (assumed approximately fixed) background. But I don’t really know mean field theory in physics so I can’t be sure the analogy is apt.