E. P. Cooper
Message
I am an ITAR US person. I do not have a secret or top secret clearance.
24
43
I don't think this works because I have a lot of comments in old posts
I'm having trouble following your overall argument, but it seems that this individual point could be addressed by making biographies more prominent in the user interface. Then you could edit your bio to clearly state your new policy (on the first or second line).
(Maybe to the point where an article with a single author would have a biography snippet at the top of the page, right after the author name.)
Presumably not exclusively. You set up a computer to regularly fetch all new articles and comments via the API (see what is sometimes called the "firehose" on bigger sites) and have it collate everything new from the users you follow. This can be done anonymously when you set it to simply download everything, if you are able to send envelopes filled with cash to particular addresses (a perfectly legitimate arrangement). This has been known for decades (see my previous citation), but most other sites do their best to prevent this type of access. Wei Dai already knows how API access to Less Wrong works, so I didn't elaborate previously.
Currently, there is (relatively) open API access that lets anyone get notified soon after you write an article or make a comment, without having to provide an E-Mail address. That seems like a non-negligible reason to keep doing what you've been doing, since the ability to anonymously and locally generate a newsletter from online sources by intentionally over-requesting data is now uncommon, even though it is useful[1].
True Names by Vernor Vinge (1980), Page 48, "If a user copied the entire board, and then searched it[...]
This random drift seems obviously caused by humans' lack of sufficient intelligence/agency/cooperation, not some sort of fundamental problem with slack?
Do you have a preferred formalism for this? UDT doesn't have a standard temporal coherence theorem by default, since even allowing some deviation from EUM (geometric means, logical induction, etc.), the UDT "prior" can be reinterpreted as a probutility measure, with only minor distortion (magnitude of distortion conjectured acceptable for humans) caused by the structure of the UDT chains provided by your mu...
Vladimir Nesov has a suggestion here about how this could be done[1]. I don't think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov's proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov's design).
...Maximax plays "defect" against an opponent with a trembling hand in PD[1] so "cooperate" can't be defended as "optimistic" here.
However, the action could be labeled "trusting" (derogatory) but not "trusting" (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since "defect" is the only way to get the best outcome (i.e. get DC), calling an agent that plays C "optimistic" is like calling someone who plays an option that perfectly bans all AI forever a "Singularitarian optimist" (even the s...
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn't work outside "fair" environments (that impression would be false). Therefore, take the following as an approximate description of what I've already written, though I will try to maintain general accuracy.
Consider a prisoner's dilemma setup, between an extremely intelligent human and a machine. The human's favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always...
I'm not sure what you meant earlier by "more than a theoretical issue" so I'll answer in two ways.
If you mean "theoretical" as in something that would need to be solved in order to use the decision theory in an aligned AI, well that's the whole thing, isn't it? Even if you want to separate that aspect out (vs. human use), you just can't. These decision theories operate by construction, particular construction, and there are no "small variations that don't philosophically matter," at least until we can solve the "decision theoretic equality" problem on abst...
since the actual calculation of that fact takes more than one observer-moment, i.e. I can't verify it all at once
As far as I can tell, this is not allowed by UDT 1.1[1].
(UDT 1.1 also needs a hard coded search order (iteration order) in order to self-cooperate properly, an additional reason why I don't think humans can run it.)
As for other agents, Wei Dai's original UDT 1.0 is logically omniscient in the sense that it will never notice its "thoughts" (such as they are) taking any noticeable amount of time. It features an "intuition module" for mathematics, ...
what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness
I'll quote Soto directly here, for future reference in some case where the Google Drive link stops working (as they sometimes do):
...In the empirical case,
could just take the algorithm that the agent is running to take decisions (for example, our messy brain synapses in the case of humans), feed it different empirical observations, and see how it reacts. There is no logical explosion, since the empirical counterfactual is perfectly consistent (just
Part of the issue is that it's pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness.
Appendix D of Martín Soto's draft report "Logically Updateless Decision-Making" (2023)[1] gives a description of the apparent metaphysical impossibility of reified logical updatelessness. As of 2023, Soto appeared to have an interest in a way of examining a particular counterlogical world that is stable under ever-increasing application of compute, though I'm not sure if this was interest dr...
AI 2040's space supplement says "[are] there any mitigating factors, e.g., the experience being necessary for a strongly positive life[...]" which could be construed as referring to mindcrime trading off against prediction quality, among other things, though barely. Is there some agreement to avoid talking about mindcrime more directly? Normally you don't have to, since the problem needs to be handled alongside malign entities and code that diverts your physical computer from the faithful execution of the abstract program/has physical side effects, e.g. ro...
My unusual claim in this area is that we had an aborted attempt to transcend the modernist representations when they proved insufficient in the face of problems in foundational math[...]
I think this would be worth working on (because of Löb) if we expected to have enough time, but plausible developments in logical induction should get within epsilon of perfect (in expectation), enough for the next few million years. Though, if you research just enough to write a document that makes the strategy semi-legible, work could be resumed if a credible long-term AI pause is implemented.
- Agree UDT is a lot better than FDT.
I'll write this in a way that's generally useful to readers since I don't know what you like and dislike about UDT, and I don't know what the term "UDT" points to in your head. Do you want to elaborate? Before, you said you liked "updateless EDT" (which UDT is not, but it is of the same theme, maybe I'll get back to this later) more than FDT because it has fewer missing components.
Note that I'm not sure I understand Bomb correctly, so please give me feedback if I get it wrong. My understanding is that the bomb going off o...
Actually, the difficult procedure in E can, by itself, lose you certain commitment races against certain opponents, if you take my short summary written there as a guide on how to think.
Note that to avoid losing a commitment race (by reaching an "infohazard") to the maximum amount UDT 1.0 allows, you must be fully updateless (base your policy on the prior alone[1]) and also have a "mathematical intuition subroutine" that is somehow optimal and doesn't (de facto) tell you too much. Intentionally making your mathematical intuition worse is an unsolved proble...
To reference, here is my list of issues that fairly sophisticated reasoners get wrong when thinking about FDT.
A: The utility function you imagine using is not gerrymandered enough. For example, it's not the decision theory's problem if you care a large amount about a small-in-prior region but forget to encode such.
B: There is not a proper attempt at writing a bounded procedure, instead the "decision theory" as imagined consists of a human trying to guess the output of a logically omniscient, unbounded procedure.
Both you and Nate Soares don't seem to have a...
Here’s the idea: you don’t really know if you are the algorithm being simulated in Newcomb’s problem or the actual person.
The history of that idea:
Garry Drescher and Eliezer Yudkowsky, in their most polished respective works, were careful to avoid positively claiming that the "agent" in Omega's prediction was the same as the actual agent, or had any experiences, and were also careful to make the reasoning work exactly the same either way.
However, this isn't what is done in many cases. Even restricting to researchers that have extremely strong claims that t...
How much worse is 1.0 than 1.1 in practical situations? UDT 1.1 has excellent theoretical properties, but to my knowledge hasn't been useful for further work so far. Non exhaustively, this seems to be because attempts to reduce the size of the outer loop and spread the computational work out over time don't work properly in 1.1, due to the use of an optimal global strategy.
Here, I'll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of "fair" scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do... (read more)