Document Title

Counterfactual mugging is the typical problem used to motivate updateless decision theory:

Omega flips a coin. If Heads, it asks you for $1. If Tails, it offers you $5 only if it predicts you would have given it the $1 had the coin come up Heads.

Counterfactuals are confusing though, so let us re-phrase it this way:

Omega flips a coin. If Heads, it offers you $(−1|H⟩+5|T⟩) (i.e. -1ifthecoincameupheads; +5 if it came up tails). If Tails, it simulates a version of you in which the coin came up Heads, and lets that simulated version make the decision as to whether to take the offer.

[Take a moment to convince yourself that these are indeed equivalent.]

This is already arguably more convincing than the original problem: if Omega comes to you and says the coin came up Heads, then you might be in the simulated world, and might actually be choosing for the outside version of you. Of course, the simulated version of you needn't care about the outside version of you, but a rational agent would self-modify to have identical versions of itself care about each other.

This "you only choose in some worlds" logic may remind you of Psy-kosh's non-anthropic problem or Conitzer 2017's Dutch book for EDT Sleeping Beauties (see also my simplified version here). Indeed, soften the problem as follows:

Omega flips a coin. If Heads, it has a 90% chance to LetYouChoose. If Tails, it has a 10% chance to LetYouChoose. If it LetsYouChoose, it offers you $(−1|H⟩+5|T⟩), and your decision in this case will be used in all circumstances.

This "all circumstances" needs to be made precise: we must ensure the simulated cases add up to the correct 90% 10% 10% 90%. E.g.

There are 10 identical copies of you, with a shared bank account. Omega flips a coin. If Heads, it gives 9 of you a GreenMarble. If Tails, it gives 1 of you a GreenMarble. For those with a GreenMarble, it offers a single payment of $(−1|H⟩+5|T⟩) to the bank account.

Now it is apparent that Psy-kosh's problem is in fact exactly the same.

There are 10 identical copies of you, with a shared bank account. Omega flips a coin. If Heads, it gives 9 of you a GreenMarble. If Tails, it gives 1 of you a GreenMarble. For those with a GreenMarble, it offers a single payment of $(6|H⟩−26|T⟩) to the bank account.

[Actually, Psy-kosh's problem is negated, but this is OK: you still have two choices; saying YES in the first case is equivalent to saying NO in Psy-kosh's case. Basically in the 90% -> 100% limit of Psy-kosh's problem, Omega says: I'll give you $1, but if you accept it then I would have murdered you if the coin had come up tails. So accepting the $1 in Psy-kosh is equivalent to not paying the $1 in Counterfactual Mugging.]


A way to beat superrational/EDT agents?

Suppose you have two identical agents with shared finances, and three rooms A1, A2, B.

Flip a fair coin.

(At each point, flip another fair coin to decide the permutation, i.e. which agent goes to which room.)

Now to each agent in either A1 or A2, make the following offer:

Guess whether the first coin-flip came up heads or tails. If you correctly guess heads, you both get 1 *  * .Ifyoucorrectlyguess *  * tails * *,youbothget * *3. No negative marking.

The agents are told which room they are in, and they know how the game works, but they are not told the results of any coin tosses, or where the other agent is, and they cannot communicate with the other agent.

...

In terms of resulting winning, if an agent chooses to precommit to always bet heads, its expected earnings are 1 * *,butifitchoosestoprecommittoalwaysbet *  * tails * *,itsexpectedearningsare * *1.50. So it should bet tails, if it wants to win.

But consider what happens when the agent actually finds itself in A1 or A2 (which are the only cases it is allowed to bet): if it finds itself in A1, it disqualifies the TT scenario, and if it finds itself in A2, it disqualifies the TH scenario. In either case, the probability of heads goes up to 2/3. So then it expects betting heads to provide an expected return of 1.33 * *,andbetting *  * tails *  * toprovideanexpectedreturnof * *1. So it bets heads.

(There are no Sleeping Beauty problems here, the probability genuinely does go up to 2/3, because new information -- the label of the room -- is introduced. BTW, I later learned this is basically equivalent to the scenario in Conitzer 2017, except it avoids talking about memory wiping or splitting people in two or anything else like that.)

What's going on? Is this actually a way to beat superrational agents, or am I missing thing? Because clearly tails is the winning strategy, but heads is what EDT tells the agent to bet.