It doesn't sound quite right to me that there are different possible cultures for any given number of echoes. I think it's more like... you memoize (compute on first use and also store for future use) what will or is likely to happen, in a conversation, as a result of saying a certain kind of thing. The thrust, or flavor, or whatever metaphor you prefer, of saying that kind of thing, starts to be associated with however the following conversation (or lack thereof) seems likely to go.
People don't actually have to be aware at all of all the levels at any one time. Precomputed results can themselves derive from other precomputed results. Someone doesn't have to be able to unpack one of these chains at all to use it. Sometimes some of the earlier judgments were actually made by someone else and the speaker is just parroting opinions he or she can't justify! (This is not necessarily a criticism. Each human does not figure everything out from scratch for himself or herself. In the good cases, I think the chain probably could be unpacked through analysis and research, if needed.)
But there remains something like the "parity" (evenness or oddness) of the process, in addition to its depth. (More depth is of course good, as long as it's accurate. It often isn't accurate, and more levels means more chances for it to diverge. I would guess this is the main reason some people (often including me) prefer lower depth - they don't expect the higher depth inferences to be sufficiently accurate to guide action. As they often aren't.) This manifests as whether we look for fault in the speaker or in the listener. This too is of course not a single value, but it's an apportionment, not a number of echoes. There is (I think) a tendency to look more towards the speaker or the listener(s) for fault (or credit, if communication goes well!), and THAT is what I think ask and guess culture are about. It ends up being something like the sum of a series in which the terms have a factor of (-1)^n.
(I agree with the overall thrust of this post that "you could just not respond!" references an action that, while available, is not free of cost such that one can simply assume that leaving a comment will consume none of the author's time and attention unless he or she wants it to.)
I am saying you do not literally have to be a cog in the machine. You have other options. The other options may sometimes be very unappealing; I don't mean to sugarcoat them.
Organizations have choices of how they relate to line employees. They can try to explain why things are done a certain way, or not. They can punish line employees for "violating policy" irrespective of why they acted that way or the consequences for the org, or not.
Organizations can change these choices (at the margin), and organizations can rise and fall because of these choices. This is, of course, very slow, and from an individual's perspective maybe rarely relevant, but it is real.
I am not saying it's reasonable for line employees to be making detailed evaluations of the total impact of particular policies. I'm saying that sometimes, line employees can see a policy-caused disaster brewing right in front of their faces. And they can prevent it by violating policy. And they should! It's good to do that! Don't throw the squirrels in the shredder!
I don't think my view is affluent, specifically, but it does come from a place where one has at least some slack, and works better in that case. As do most other things, IMO.
(I think what you say is probably an important part of how we end up with the dynamics we do at the line employee level. That wasn't what I was trying to talk about, and I don't think it changes my conclusions, but maybe I'm wrong; do you think it does?)
I have trouble understanding what's going on in people's heads when they choose to follow policy when that's visibly going to lead to horrific consequences that no one wants. Who would punish them for failing to comply with the policy in such cases? Or do people think of "violating policy" as somehow bad in itself, irrespective of consequences?
Of course, those are only a small minority of relevant cases. Often distrust of individual discretion is explicitly on the mind of those setting policies. So, rather than just publishing a policy, they may choose to give someone the job of enforcing it, and evaluate that person by policy compliance levels (whether or not complying made sense in any particular case); or they may try to make the policy self-enforcing (e.g., put things behind a locked door and tightly control who has the key).
And usually the consequences look nowhere close to horrific. "Inconvenient" is probably the right word, most of the time. Although very policy-driven organizations seem to have a way of building miserable experiences out of parts any one of which might be best described as inconvenient.
I'm not sure I agree who's good and who's bad in the gate attendant scenario. Surely getting angry at the gate attendant is unlikely to accomplish anything, but if (for now; maybe not much longer, unfortunately) organizations need humans to carry out their policies, the humans don't have to do that. They can violate the policy and hope they don't get fired; or they can just quit. The passenger can tell them that. If they're unable to listen to and consider the argument that they don't have to participate in enforcing the policy, I guess at that point they're pretty much NPCs.
I don't know whether we know anything about how to teach this, other than just telling (and showing, if the opportunity arises), or about what works and what doesn't, but I think this is also what I'd consider the most important goal for education to pursue. I definitely intend to tell my kids, as strongly as possible, "You always can and should ignore the rules to do the right thing, no matter what situation you're in, no matter what anyone tells you. You have to know what the right thing is, and that can be very hard, and good rules will help you figure out what the right thing is much better than you could on your own; but ultimately, it's up to you. There is nothing that can force you to do something you know is wrong."
I hadn't noticed that there'd be any reason for people to claim Claude 3.7 Sonnet was "misaligned", even though I use it frequently and have seen some versions of the behavior in question. It seems to me like... it's often trying to find the "easy way" to do whatever it's trying to do. When it decides something is "hard", it backs off from that line of attack. It backs off when it decides a line of attack is wrong, too. Actually, I think "hard" might be a kind of wrong in its ontology of reasoning steps.
This is a reasoning strategy that needs to be applied carefully. Sometimes it works; one really should use the easy way rather than the hard way, if the easy way works and is easier. But sometimes the hard part is the core of the problem and one needs to just tackle it. I've been thinking of 3.7's failure to tackle the hard part as a lack of in-practice capabilities, specifically the capability to notice "hey, this time I really do need to do it the hard way to do what the user asked" and just attempt the hard way.
Having read this post, I can see the other side of the coin. 3.7's RL probably heavily incentivizes it to produce an answer / solution / whatever the user wanted done. Or at least something that appears to be what the user wanted, as far as it can tell. Such as (in a fairly extreme case) hard coding to "pass" unit tests.
I wouldn't read too much into deceiving or lying to cover up in this case. That's what practically any human who had chosen to clearly cheat would do in the same situation, at least until confronted. The decision to cheat in the first place is straightforwardly misaligned though. But I still can't help thinking it's downstream of a capabilities failure, and this particular kind of misalignment will naturally disappear once the model is smart enough to just do the thing, instead. (Which is not, of course, to say we won't see other kinds of misalignment, or that those won't be even more problematic.)
That's possible, but what does the population distribution of [how much of their time people spend reading books] look like? I bet it hasn't changed nearly as much as overall reading minutes per capita has (even decline in book-reading seems possible, though of course greater leisure and wealth, larger quantity of cheaply and conveniently available books, etc. cut strongly the other way), and I bet the huge pile of written language over here has large effects on the much smaller (but older) pile of written language over there.
(How hard to understand was that sentence? Since that's what this article is about, anyway, and I'm genuinely curious. I could easily have rewritten it into multiple sentences, but that didn't appear to me to improve its comprehensibility.)
Edited to add: on review of the thread, you seem to have already made the same point about book-reading commanding attention because book-readers choose to read books, in fact to take it as ground truth. I'm not so confident in that (I'm not saying it's false, I really don't know), but the version of my argument that makes sense under that hypothesis would crux on books being an insufficiently distinct use of language to not be strongly influenced, either through [author preference and familiarity] or through [author's guesses or beliefs about [reader preference and familiarity]], by other uses of language.
I agree that the average reader is probably smarter in a general sense, but they also have FAR more things competing for their attention. Thus the amount of intelligence available for reading and understanding any given sentence, specifically, may be lower in the modern environment.
Question marks and exclamation points are dots with an extra bit. Ellipses may be multiple dots, but also indicate an uncertain end to the sentence. (Formal usage distinguishes "..." for ellipses in arbitrary position and "...." for ellipses coming after a full stop, but the latter is rarely seen in any but academic writing, and I would guess even many academics don't notice the difference these days.)
I read a bunch of its "thinking" and it gets SO close to solving it after the second message, but it miscounts the number of [] in the text provided for 19. Repeatedly. While quoting it verbatim. (I assume it foolishly "trusts" the first time it counted.) And based on its miscount, thinks that should be the representation for 23, instead. And thus rules out (a theory starting to point towards) the corrects answer.
I think this may at least be evidence that having anything unhelpful in context, even (maybe especially!) if self-generated, can be really harmful to model capabilities. I still think it's pretty interesting.
I have very mixed feelings about this comment. It was a good story (just read it, and wouldn't have done so without this comment) but I really don't see what it has to do with this LW post.
I would guess this is about "getting the right things into context", not "being able to usefully process what is in context". (AI already seems pretty good at the latter, for a broad though not universal set of tasks.)