I think that only gets you EDT.
Relevant to my point is that I don't see why PPO'd always get a tickle defense, since it doesn't need one to work AFAICT.
Yes it can only work consistently if the reward-model isn't much dumber than the agent. Is that uncommon for some reason? I was assuming you'd sometimes use the same model for both, maybe even some clever trick to be able to use the same instance.
In that example the reward-model gets to see what
Does PPO generally work that way? I'd've guessed it doesn't always get a tickle defense, I'm not an engineer. If it does then I guess this stuff doesn't matter for PPO.
I googled it and found this paper from @Caspar Oesterheld. IIUCT some common algorithms create EDT, that's what I meant it as an example of.
Thanks! You understand.
Follow-up argument. By selecting from the class of optimal decision theories, we are applying infinite selection-pressure on the set of all decision theories. Understandably this produces perverse incentives toward sabotaging other decision theories (the same way GRPO incentivises models to sabotage other instances of themselves). But in this context we really do want to choose a decision theory that always does at least as well as any other, otherwise we're leaving utility on the table.
I think GRPO definitely creates CDT in-the-limit (assuming no scheming mesaoptimizers like what you describe). PPO creates EDT though.
I think PPO creates EDT in-the-limit.
Yes of course rewarding models for doing better than their peers will cause them to defect against their peers (and in fact try to hurt them).
How'd this go?
That proof doesn't make sense to me. LDT gets $1 against $9-rock, and no other agent can get more than that against $9-rock. So where's the suboptimality?
Our universe happens to be continuous, so by the IVT any physical implementation of that decision problem will allow XDT a reflectively-consistent mixed-strategy. If it has enough compute, it can approximate that strategy and in-expectation receive epsilon less money than the best agent will.
@Joar Skalse Is this right, or salvagable? If so, does it matter?
When weird ideas are discovered, or very weird things happen, generally it's found that the set of cases includes very few non-edges. ASI is very important and very weird, so things will go wrong even aside from its being adversarial.
"tendency to 'bite bullets' or accepting implications that are highly counterintuitive to others or even to themselves, instead of adopting more uncertainty"
If people changed their minds easily, clearly that would demonstrate overconfidence. So it seems like changing their minds too seldom shouldn't do the same.
"and much of my actual eyebrow-raising at this space of hypotheses comes from the way that i expect the end result to be quite sensitive to the processes of reflection"
Says the CEV guy
The main point of the "Coherent" in CEV is that, since regular old Extrapolated Volition might be very sensitive to how extrapolation and volition are defined, probably as much as feasible we want to put off final decisions there rather than letting the programmers make them.
If that's surprising, probably it's because CEV is a poorly-chosen name for the concept.
Edit: Nate als...
Huh, I remembered the original post as being intended to apply specifically when subagents are bad at coordinating. Feels bad to have known better than someone and assumed he already knew what I did (or else failed to tell him for some reason?).
I think the post is mostly about coolness relative to other kinds of status. I feel fine believing that John has a greater propensity to coolness than to other kinds of status.
The advice is mostly meant to reduce pressure from conformism, and only narrowly to improve any social skills. As far as I can tell.
It was presented by Bean, not Ender. At least according to the linked post.
It's easy to estimate cos(floor(pi * 3^^^3)) to within 30%. Try it!
This is true for most concepts.
It's an especially useful frame for free will, moreso than for temperature I think. Free will's subjectivity is relevant because (unlike for temperature) humans can in fact outgrow its use in some contexts.
If you have the payoff matrix yes it's CDT.
You could have each instance of the agent see a random seed before being predicted, such that they have different probabilities of two-boxing. ) back to the that used to be conventional for REINFORCE, you maybe even get game-level updatelessness in a multi-turn game, I'm having trouble thinking about it.)
(In that case if you change your rightmost lower bound (