Showing posts with label Reverse Monty Hall problem. Show all posts
Showing posts with label Reverse Monty Hall problem. Show all posts

Tuesday, 26 September 2017

cnearing's argument, subjective probability and the whole Reverse Monty debacle

Over at Craig-Land, in his thread on what constitutes a "good argument", a forum member called cnearing wrote this:

I would point out that I take a strictly subjectivist approach to probability.  Probability is not an objective feature of reality (probably--set aside potential quantum weirdness for the moment) but rather a representation of the uncertainty one has about the way things behave.

Probability comes from models.  Models are built through inference from experience.  Both are subjective.

This reminded me of the terrible trouble I got myself into with respect to the "Reverse Monty Hall Problem".  I went into it with an assumption which, for one reason or another, had become implicit rather than explicit and which then became forgotten, hidden and/or overlooked.

There are a lot of articles connected with the one I've provided a link to, so I'll try to distil it down to essentials.

If we don't know, and have no way to know, the probability of a particular claim or premise in terms of it being true or false, we are required by the principle of indifference to assign it a notional probability of 1/2.  If we know more, for example that there are more options, say:

A is true, B and C are false.
B is true, A and C are false.
C is true, A and B are false.
One of A, B and C is true.

Then, knowing nothing else, we are required to assign a notional probability of 1/3 to A, B and C.  We don't know anything that makes A more or less likely than the others.

Say then, that we have an urn in which there is an arbitrarily large number of balls that are identical in size and that is all we know.  We draw out all but one of them (say 999 of them), and they are all white.

What is the likelihood that the last one is also white?  If we know nothing else, then the answer is one in a thousand, the same likelihood of a single non-white ball being placed in any specific position in a sequence of one thousand extractions.

What was the likelihood, after having drawn 99 balls, that the 100th was non-white?  The same logic applies and it was one in one hundred.  As more white balls were drawn, the likelihood of a non-white ball being drawn went down.

Now, compare this to the likelihood of drawing a non-white ball as assigned by someone who has one more piece of information, the knowledge that the balls were initially drawn from an enormous barrel in which there were 900,000 white balls and 100,000 black balls.

Unlike us, who had no additional information and could only work on the basis of what balls we had drawn out, this other person will know that there is an increasing likelihood of drawing a black ball after each white ball is removed.  In fact the likelihood of the 1000th ball being non-white after a sequence of 999 white balls is very slightly higher than one in ten.  The likelihood of drawing 999 non-white balls in a row is extremely low, but that is immaterial, since we are only looking at the likelihood associated with the next draw once this extremely unlikely scenario has already played out.

We can fiddle with the figures to make it more explicit.  Say we only know about 999,999 balls that we've drawn, over a period of a couple of boring days.  All of them are white.  We have to say that the likelihood of the next ball being non-white is one in a million.

But if our more knowledgeable friend knows also that there is one black ball in the barrel, then she will have to say that the likelihood of the next ball being non-white is one in one, 100%.

---

My point here is that we have evidence, and we might also have assumptions.  The assumptions that we make about the distribution of balls in the urn will change our assessment of the likelihood of a non-white ball being drawn.  If we characterise our (potentially false) assumptions as "knowledge" - as theists are often wont to do - we will consistently misjudge the likelihood of our premises (and subsequent conclusions) being true.

Add to this the possibilities that we don't consider (i.e. thinking only of A being true or false, rather than factoring in other possibilities like B and C), then we end up with very little likelihood of reaching reliable conclusions.

Sunday, 22 November 2015

Monty Chooses

I want to revisit Monty Hall again, for reasons that hopefully will become apparent.

Suppose there was a slight variation to the standard Monty Hall scenario.  Consider the following:

Monty Hall is the compere on a game show.  In one special round of that game show there is the familiar set-up.  There are three doors, behind one of which is a car and the other two doors each hide a goat (placed randomly).  The contestant is invited to select a door, with the idea that if the car is behind her car, she will win it.  The contestant selects a door (we’ll call this “door X”, meaning “the door that the contestant chose”).

The main observable difference is that Monty now selects a door (we’ll call this “door Y”, meaning “the door that Monty chose”).  He goes up to that door and looks through a peephole.  Then Monty walks to the other door (which we’ll call “door Z”, meaning “the door that neither chose”), and opens that door, revealing a goat.

Given that we know nothing else, can we work out the likelihood that the contestant will benefit by swapping to door Y, if given the opportunity?

It looks very similar to the standard Monty Hall problem, but is the difference significant?

It’s tempting to assume that Monty will never reveal a goat, so he would have opened door Y if there was no goat there, but he didn’t.  He walked to door Z and opened it, revealing a goat.  Why would he have done that?  The hints are that if he never reveals a goat, then he must have known that the goat was not behind door Z, and the fact that he peeped at what was behind door Y indicates that he hasn’t been informed as to the distribution of goats and car, so the car must be behind door Y.  Therefore, the contestant should, in order to win the car (with an apparent likelihood of 100%), swap from door X to door Y.

Now consider instead, holding all other factors equal, a slightly different scenario in which Monty peeps behind his door, door Y, and then opens it.  What is the likelihood that the contestant will benefit from swapping to door Z?  Is it the same as Monty Falls, in which he accidentally opens the first door he passes and it just happens to reveal a goat, or does this situation mirror the standard Monty Hall problem?

To work it out, we have to remember that nothing was stated about Monty’s intentions nor about what he knows and doesn’t know – and it was specifically stated that we know nothing else.  We weren’t even told that the contestant would get the chance to swap, which she wouldn’t have if Monty had revealed a car (this possibility was not excluded, if you thought it was, this was just an assumption on your part).

Sticking with the scenario in which Monty opens door Y, all the contestant knows is that a goat is behind door Y and that Monty knew this when he opened the door.  She does not know how Monty selected the door to peep through, but we are tempted to assume it was random (or that it doesn’t matter because the distribution of car and goats is random).

If we make a different assumption, we get a different likelihood that the contestant would benefit from swapping.  Say that Monty is restricted from revealing the car and he knows where it is.  This does not prevent him from selecting a door to peep through at random.  There is a 1/3 chance that the contestant’s door hides the car, in which case he can choose to open either door Y or door Z.  There is a 1/3 chance that the door he peeps through hides the car, in which case he must open the other door (we’re assuming that he can’t open the contestant’s door).  There is a 1/3 chance that he has no choice and must open the door he peeps through (again assuming that he can’t open the contestant’s door).

Monty’s subterfuge

A general assumption in the standard Monty Hall problem is that, when given a choice as to which door to open, Monty will simply select one at random.  This is predicated on the assumption true that Monty doesn’t care whether the car is won.  However, if he does care and he gets the opportunity to peep behind a door, despite knowing the distribution of car and goats, then he has motivation to make thoughtful rather than random selections.  In this case does it matter which decision he makes and what decision should he make to minimise the likelihood of the contestant winning the car?

I suspect that it does matter and that his best option is to preferentially open door Z.  The reason for this is that the contestant does not know what is known to Monty and is likely (more on the basis of psychology than of statistics) to assume that there is a greater likelihood that Monty is forced to open door Z than that he had done so by choice.  (This would be the case if he chose at random when given the opportunity and also if he didn’t know the distribution of car and goats.)  Therefore she will use the logic laid out above to conclude that the car is behind door Y.  But Monty will choose door Z more frequently than the contestant thinks, 2/3 of the time – when the car is behind door Y and when the car is behind door Z.  So, she’ll only win the car 50% of the time by swapping to door Y.

If, on the other hand, Monty opens door Y after peeping behind it, then contestant might easily conclude that she is in a Monty Falls type situation (based on the assumption that Monty selects the door at random and, lo, it doesn’t have the car, so he opens it) and that therefore she doesn’t benefit statistically from swapping (or indeed from staying).  There are psychological factors that make us prefer things we possess, so if the contestant thinks that it’s 50-50 as to where the car is, she’s more likely to stick with her choice, meaning she’ll lose.  If she leaves the choice up to chance, tossing a fair coin to decide, then there’ll be a 1/2 chance that she’ll swap, and thus a 1/2 chance of winning.  If she simply ignores the extra information implicit in this scenario, Monty’s peeping behind door Y, she’ll win.

By engaging in this subterfuge, Monty can possibly minimise the likelihood that the car will be won, but only if it remains a secret.  If the contestant becomes aware of Monty’s subterfuge, she’ll know that if door Y is opened then the car is definitely behind door Z, but if door Z is opened then it’s 50-50 as to where the car is.  Her best strategy would then default to that of the standard Monty Hall, always swap.  In other words, she should ignore an element of information available to her to maximise her chances of winning.

Helpful Monty

Monty could, however, have a different set of intentions.  He might desperately want to give away a car, but be obliged to remain within the rules of the game (an equivalent to this appears to be a standard assumption within intelligent design theory).  If so, he merely swaps his strategy to encourage a consideration on the part of the contestant that would maximise her chance of winning – by preferentially opening door Y, only opening door Z if the car is behind door Y.  In this case, the contestant will have the same likelihood of winning whether she pays attention to Monty’s peeping or not.  It also doesn’t matter whether she believes that Monty has no idea where the car is, or if she thinks he’s trying to help.  In this case, she can ignore the information or try to make use of it, it doesn’t really matter either way.

There’s Something about Mary

Alternatively, Monty might not care about the car at all and merely prefers to reveal Mary (one of the goats).  If so, he’ll open the door Z if Mary isn’t behind door Y when he peeps, even if that might reveal the car, because it maximises his chance of revealing Mary.  If the contestant knows this to be the case, then she would know that there’s a benefit in swapping to door Y, if door Z was opened, but it’s 50-50 if he opens door Y.  There’s a 1/3 chance of Mary being behind door Y.  There are two equally likely distributions in which Mary could be behind door Y, one which has the car behind door Z, to which the contestant might consider swapping.  There is a 2/3 chance that Mary will not be behind door Y in which case Monty will open door Z.  There is a 1/2 chance that he did so after seeing the other goat and if so, there is a 1/2 chance that when he opens door Z, he’ll reveal the car – nullifying the game.  This gives the contestant a 2/3 chance of winning if Monty opens door Z and she swaps.  The overall likelihood of winning as a result of swapping is 3/5 rather than the standard 2/3 because of the 1/6 chance that Monty will reveal the car.

However, yet again, the contestant can obtain the same result merely by ignoring Monty when he peeps behind the door.

Now for the key question: So frigging what?

Whoa, that was unnecessarily aggressive!  The point is that Monty is an intelligent agent, with a psychology that the contestant must intuit.  The same sort of consideration applies in Hawthorne’s prisoners, only in that case the (ultimate) psychology being considered was that of a putative god.

What I’ve tried to show here is that the intent of the intelligent agent can matter sometimes, when working out the likelihood of desired result, but it doesn’t always matter – sometimes it has little or no apparent influence at all.  The trouble is that it’s not exactly obvious as to what matters.  It matters if Monty is committed to not revealing the car, and it could matter if we are persuaded that he doesn’t know where the car is, but we can maximise the likelihood of a good outcome by simply ignoring his intent.  Similarly, for the purposes of working out whether Miss Justice is responsible for a release decision, it matters how she would choose a prisoner to release and whether their location matters, but so long as location doesn’t matter it won’t really matter much what her criteria actually are – meaning that our ability to detect her influence will not vary if she prefers to release the shortest innocent prisoner rather than one who is most worthy of clemency.  An implication that we can draw from this is while it certainly does matter what an intelligent agent knows, what they intend is somewhat less likely to matter.

This has some bearing, I think, on the fine tuning argument for the existence of some sort of god.  Fundamentally, the argument derives from an observation that the universe has a suite of constants and initial conditions that could not be varied much without making life impossible.  A designer is then posited as the mechanism by which the values for these constants and initial conditions were selected – with the assumption that the designer intends that (intelligent) life should manifest.

While there is a danger going from the specific to the general, let alone from one specific to another specific, what the consideration above (together with Weisberg’s prisoners) tells us is that we should at the very least be careful when basing an argument on an assumed intent.  The fact that theists base their best argument * on such foundations should give them pause.

---

* Admittedly this is the theists’ best argument according to Christopher Hitchens.  Perhaps theists themselves think that they have a better argument.

Wednesday, 7 October 2015

(My) Ignorance Behind "Marilyn Gets My Goat"

At the end of Marilyn Gets My Goat, I wrote:

As mentioned above, this argument is wrong. Precisely which bit is wrong is a little vexed. I originally thought the biggest issue nestles in Q10 and Q16. I still think it does, but many below argue that the issue is in Q15.  That might come down to a question of interpretation, but eventually I will put it into words, hopefully without sparking huge controversy this time.

I was initially quite happy with the argument that I had presented (especially in combination with a later clarification in Marilyn's Six Games).  Even now, looking at it with the knowledge that it is wrong somehow, I still find it faintly convincing.

I suspect that this has to do with ignorance.  Clearly I was overwhelmed with ignorance when I started off this merry chase, but it's not that sort of ignorance that I mean.  I mean more the type of ignorance that I mentioned in The Whole Reverse Monty Debacle:

… I had in my mind a scenario in which the volunteer (now transforming into a contestant) knew nothing, other than the nature of the revealed ball (now transforming into a goat).  During the transformation process, I forgot all about that ignorance and began trying to apply my thinking (in the presence of ignorance) to the Monty Hall Problem (in which there is less ignorance).

The problem, I believe, is associated with the doors.  The doors represent ignorance, because the contestant doesn't know what is behind the doors.  But clearly they don't consistently represent ignorance, because Monty and Holly apparently do know what is behind the doors.

To try to get at the heart of the problem, I want to go over the step-wise scenario again, after removing the doors and everything behind them.  Imagine instead that we are left only with Holly Mant and Marilyn, who is given a three-sided die (yes, they exist) and a coin.  On the table between them are three trays each with six tokens.


These trays represent the "expected value" of each of the doors that we've removed but will still think about as concepts.  We've assumed a random distribution of goats (M and A) and car (C), so there is a 1/3 chance of M behind each door, a 1/3 chance of A and a 1/3 chance of C.  This aligns with Marilyn's answer to Q1.

The arrangement of the tokens indicates the possible distributions of the car and goats, as given by Marilyn in response to Q2, and simultaneously show that the likelihood of each distribution is 1/6, as per the answer to Q3.

Marilyn is now encouraged to roll her three sided die (helpfully labelled NOT RED, NOT WHITE and NOT GREEN - as per the response to Q4).  Following the scenario in Marilyn Gets My Goat this would mean she rolls NOT WHITE and selects Red and Green.  The likelihood of selecting these two doors is 1/3, presuming that this is a fair die, and aligns with the answer to Q5.

So far so good.  Let's reorganise the trays:


Q6 is a gimme, since it's effectively the answer to "given X, what is the likelihood of X?"  And we can see the answers to Q7 and Q8 laid out in the selected trays - each distribution has a likelihood of 1/6.

Now it gets trickier.  Marilyn is asked what the likelihood is that the red door will be selected (which Holly won't do if there's a car behind it).  That can be represented like this:


The result is that there is a prior likelihood of 1/2 that the red door will be selected (1/12+1/6+1/6+1/12 = 6/12 = 1/2) - as per Marilyn's answer to Q9.

Marilyn is then asked Q10: "What is the likelihood that the Red door would have to be opened, in accordance with the rules, on the basis that that the car was behind the Green door?"

I now see that this question has a problem, which is reassuring on one level, because I did identify it as possibly harbouring an issue, and annoying on another, because I was trying to eliminate confusion and an ambivalent question like this will merely sow confusion.  There are two answers to the question, depending on your interpretation.

One interpretation leaves the question precisely as it is.  There are two scenarios in which the red door must be opened because there is a car behind the green door – AC and MC.  This means that the likelihood of the the red door being opened because of the location of the car is 1/3 (1/6+1/6 = 1/3).  There is also a 1/3 chance that the green door would be opened, because of the location of the car, a 1/3 chance that whichever door was opened was opened on the basis of a random selection.

This is not the interpretation that I (or indeed anyone, so far as I can tell) used.  We assumed that the question really meant this: “Given that the Red door was opened, what is the likelihood that the Red door had to be opened, in accordance with the rules, on the basis that that the car was behind the Green door?”

To work this out, we have to divide the likelihood that the car was behind the green door AND the red door was opened by the likelihood that the red door was opened, so (1/6+1/6) / (1/6+1/6+1/12+1/12) = (1/3) / (1/2) = 2/3.  Which was the answer given to Q10.

Note that this is also precisely the time when my ignorance (as described in The Whole Reverse Monty Debacle) began to really screw things up.  I had in my mind, at least at the very beginning, the idea that the contestant didn’t actually know the basis on which Holly would open the door.  Of course I should have known, since I was drawing parallels to the Monty Hall Problem, but I still had the vestiges of a random selection paradigm in my head.

So, just for interest’s sake, let’s look at a slightly different question: “If the door is opened by Holly Mant on the basis of a coin toss, what is the likelihood that the Red door will be opened?”

We have to go back to Q9 though, or rather Q9B, because the distribution of expected value is different:


The result, again, is that there is a prior likelihood of 1/2 that the red door will be selected, but with a different calculation (1/12+1/12+1/12+1/12+1/12+1/12 = 6/12 = 1/2).

In this scenario, with a random selection based on the toss of a fair coin, there's little point in asking "What is the likelihood that the Red door would have to be opened, in accordance with the rules, on the basis that that the car was behind the Green door?" – because the location of the car is not a factor in Holly’s decision.  Note: I am aware that this is no longer parallel to the Monty Hall Problem.

We can meaningfully ask what appears to be the final question: "Given that the Red door was opened, what is the likelihood that the car is behind the Green door?"   It's not quite the final question though, because we haven't yet specified what was behind the red door, we've only said that it was opened and we've not eliminated the possibility that the car would be revealed.  Given that, we can say that the answer to what becomes Q10B is given by the likelihood of the car being behind the green door divided by the likelihood of the red door being opened: (1/12+1/12) / (1/2) = 1/3.  Alternatively, you can look at the image and see that in 1/3 of the distributions, the car is behind the green door.  This is the same as the answer to Q10, as asked.

It's also the answer to a question about a prior likelihood: "What is the likelihood that the Red door will be opened AND the car is behind the Green door?"  Remember that we are now assuming that Holly opens the door on the toss of a coin.  There will still be the six possible outcomes illustrated, and in two of them the car is behind the green door.

The difference between these two scenarios is knowledge or ignorance on the part of Holly (or, more strictly, the ability of Holly to act on any knowledge she might have).  If she knows where the car is (and can act on that basis), she skews the outcome.  If she doesn't know, and opens a door randomly, she doesn't.  Note also that Q10, as asked, is a question about a prior likelihood, in other words a likelihood calculated in the presence of ignorance.  Therefore, it should be no surprise that the results align.

Let's move on.  The next step in Marilyn Gets My Goat was to reveal what was behind the red door - and it was Mary the Goat.

Q11 and Q12 were both gimmes, although Q12 might have been a bit confusing to someone who isn't great at English.

Once Mary has been revealed to be behind the red door, the trays could look like this:


This, which graphically represents the answer to Q13, is sort of what I had in mind when I was making my argument, especially in Marilyn's Six Games.  This was of course wrong.  What I had originally thought of could have more accurately been represented like this:


But the scenario I was actually discussing was pretty much this:


Q14 was merely harking back to Q6, with the same correct response.

The general consensus was that the answer to Q15 was where I truly messed up.  I'd agree if I had asked a slightly different question to the one I did: "How likely are each of these distributions now?"  The thing is, they actually are equally likely, but Mathematician and ChalkboardCowboy were leaping ahead to provide an answer to a completely different question (namely: "Given that Mary has been revealed to be behind the Red door, is it equally likely that the car is behind the green door as not?" to which the answer is no - I now agree with that answer, although I didn't at the time).

The subtlety that we were all failing to address/make clear is that yes, the two distributions are equally likely, but every single time the first distribution comes up Holly will open the red door to reveal Mary while she will reveal Mary only half of the time in the second distribution.

---

For the sake of clarity, I'll interpret the image above.  Once the red door is opened, revealing Mary (and revealing that this is a "Red Mary game"), there are two possible distributions - either the car is behind the green door, or Ava (the other goat) is.  These are the two games that Marilyn had to consider, and each of these games is actually equally likely (thus the first 1/2 used in the equations to the left and right).

However, in every instance of the first Red Mary game, Holly is obliged to reveal Mary (thus the 1/1 against Mary and the 0/1 against the car).  Therefore, the likelihood of a game in which the car is behind the green door AND Mary is revealed to be behind the red door (a Green Car Red Mary game) is 1/2*1/1 = 1/2.

In an instance of the second of these Red Mary games, Holly may open either the red door or the green door, and will toss a coin to choose which (thus the 1/2 each against Mary and Ava).  The likelihood of a game in which Ava is behind the green door AND Mary is revealed to be behind the red door (a Green Ava Red Mary game) is 1/2*1/2 = 1/4. 

In total, there is an "expected value" of 3/4 for the red door, when we are limited to Red Mary games - which is to say that if you run a sufficiently large number of Red Mary games, the red door will be opened in 3 out of every 4 instances.

The likelihood of the car being behind the green door, if the red door has been opened to reveal Mary is the expected value of Green Car Red Mary games divided by the combined expected value of both Red Mary games, so (1/2) / (3/4) = 2/3.  This is the standard answer, and the one I should have arrived at, but of course I didn't.

---

Now to the other question where I thought I had screwed up, Q16.  Oh my!

I did screw it up.

There's nothing better when trying to confuse yourself (and potentially others) than asking a question that contains a false assumption.  I asked "What about the likelihood that the producer would instruct Holly to open the Red door being 2/3?"

This relates back to Q9, which I've already addressed.  The prior likelihood that the red door would be opened was actually calculated to be 1/2, not 2/3.  The 2/3 result was arrived at in response to Q10, which asked a more specific question.

Then I answered a totally different question, effectively "Given that the Red door has been opened revealing Mary, what is the likelihood that Mary is behind the Red door and that the Red door has been opened?"  I gave the right answer to this question, 1/1.

Finally, I raced onto provide the answer that applied to a scenario that not even I knew I had in my head.

Let's look at that one more time before we leave this all behind us forever:


This represents a situation in which either Holly has no knowledge of what lies behind the closed doors or is unable to use that knowledge, or Marilyn has no idea as to what is going on (see The Whole Reverse Monty Debacle) and makes the assumption of a random selection.  In other words, an ignorance rich scenario.

In such a scenario, the likelihood of a car being behind the green door, given that the red door has been opened to reveal Mary is 1/2 ((1/4) / (1/2) = 1/2).

If I had figured all this out first time around, I would have saved myself and some other, surprisingly tolerant people a lot of nausea.
---
If I am still wrong, I am not sure that I want to know ...

Monday, 17 August 2015

The Whole Reverse Monty Debacle

I promised that I would try to review just how I managed to convince myself that I was right with the reverse Monty Hall problem (when I was very clearly wrong).  I have the benefit of distance with which to assess what happened (because I am now looking at what a completely different person did, you know, “me” but the “me” of a few months ago, who is totally different to the “me” today).

My focus here will not be on the results of my error, but rather trying to understand how I made the error – and why I found the error convincing.  (There are quite a few posts on the results already, together with some comments to those posts as well as quite a bit of activity over at reddit, largely in the “bad mathematics” area.  Admitting I was wrong has apparently still not appeased the masses.  Oh well.  Never mind.)

I’m aware that my memory may not be perfect, partly because humans forget things and partly because humans tend to edit the past to favour themselves, so this review may not be entirely “the truth”, but I will try my best to be objective.

It all started out with me thinking about inductive logic, versus deductive logic and, sort of, about black swans.  The idea was that if you spend your whole life seeing nothing but white swans (as Europeans used to), then you’d reason that if you were presented with a swan in a box, it would be reasonable to assume that it is white.  (Sherlock Holmes might have “deduced” that the very fact that the swan was presented in a box indicates that the giver was hiding some dark secret.  However, the amazing accuracy of many of Holmes’ “deductions” relies entirely on the author being in control of the world in which Holmes lives – in the real world they would be little better than informed guesses and are better categorised as “intuitions”.  There are a host of reasons why a person might present you with a swan in a box, other than to hide the possible fact that it is a black swan – they might hiding the fact that it’s not a swan at all but rather just a particularly ugly duck.)

Anyways … I came up with a scenario in which a volunteer removes, one by one, balls from an urn.  If after 999,999 white balls have been removed, I stop them and ask them the likelihood that the next ball is not white.  In the absence of any other information, the volunteer should respond that there is a 1/1,000,000 chance that the next ball is not white.  Here’s my logic:

I’ve constrained the exercise to the removal of one million balls from the urn.  Effectively, what I’ve asked the volunteer is “what is the likelihood that, out of a draw of one million balls, the only non-white ball will be in the one millionth position?”  The other, more likely outcome, from the volunteer’s point of view, is that the last ball will also be white.  This is because the volunteer is working with an absence of any other information.  There may be more balls in the urn, of which an unknown number are non-white, or the one millionth ball might be the last ball and I, as the experimenter, might know that this last ball must in fact be black.  It’s this last scenario that I had in mind.

You see, as the volunteer keeps removing white ball after white ball, she is rightfully becoming more and more convinced that the next ball will be white, following the logic of “what is the likelihood that, out a draw of X balls, the only non-white ball will be at the Xth position?”  As X increases, after X-1 white balls have been drawn, it appears more unlikely that the next ball will be non-white.  I, on the other hand, have more information than the volunteer and am aware that the likelihood of the next ball being non-white is actually increasing.  This reaches a crescendo at the 999,999th ball at which point, I know the likelihood of the next ball being non-white is 100% while the volunteer will believe that there is only a 0.0001% chance of it not being white.

It seems to me that this is a justified conclusion by the volunteer under the circumstances, but then I also know that it is wildly inaccurate.

So, I tried to rein the numbers in a bit, all the way back to two.  Say that my volunteer reaches into an urn and removes a white ball.  In the absence of any other information, what is the likelihood that the next ball selected will be white?  It’s 50%, but this also seemed strange to me – after all, we are not just talking about the possibility of white and black balls here, there could be a huge array of hues, colours and patterns available.  Perhaps it seems reasonable to think about selecting a second pure white ball, since that’s a simple enough decoration, but if we think of a scenario in which the volunteer extracts a ball with light violet stripes and puce dots on a beige background with orange swirls, it seems somewhat less certain that there will be a 50% chance that the second ball will be of the same type.

This sort of gets us to where I was initially headed.  We tend not to notice when bland things happen (like selecting a plain old white ball) and are confounded when strange things happen (like removing our tutti-frutti themed ball), it messes with our intuitions.  The likelihood of selecting that strangely patterned ball seems remote, but it won’t be if it’s also a common pattern.  If my volunteer keeps dipping her hand into the urn and removing similarly patterned balls, then the amazement of that first selection will fade and eventually her intuitions will shift to match what would be expected with an unbroken series of plain white balls.  At that point, she may reflect that back when she held only one of these balls in her hand, the likelihood (at the time) of the next one being the same was also 50%, as it would have if it had been white.

(Note that with the information to hand, my volunteer will have to reassess the post facto likelihood of the second ball being tutti-frutti themed as being higher than 50% – the exact figure depends on the number of balls available and how many balls have been selected so far.  The 50% figure is based on maximal ignorance, and is very rubbery.  In the absence of knowledge about the distribution of balls, the volunteer should perhaps surmise only that the second ball would either be the same as the first or it wouldn't.  One of two options, with no information as to the likelihood of either option.  She would not reach the 50% conclusion if she had made the subjective assessment that the first ball was inherently unusual, but that would be another calculation: "what is the likelihood that when two balls are selected at random from an urn, that both those balls will be unusual and the same?"  Without an objective assessment of how unusual the balls are, no figure could be reached, but she could just say "not very likely at all".)

Compare this to an argument often run by apologetic theists: the probability of the universe being just the way it is such that it supports intelligent life (more specifically humans) is so remote that it is therefore inconceivable that the universe arose by chance, therefore god.  We are in the same position as my feckless volunteer, after her first selection, metaphorically holding an apparently impossible ball in our hand and being stunned and amazed by it.  However, we are not able to draw from the urn again to get a better idea about how likely it really is that such a ball should be in our hand.  So, while in the absence of any other information it may seem unlikely that our universe should be so apparently finely tuned, we simply don’t know what the real likelihood is.

I then devised a scenario to test this challenge to our intuitions, which I wrote up (at the second attempt) at Two Balls, One Urn, Revisited.  The whole idea of this scenario was to trigger the logic of “what is the likelihood that, out a draw of X balls, the only non-white ball will be at the Xth position?” in which my volunteer might say 50% where it can be shown that it’s not, it’s a wildly different figure.  In this article, I had a barrel with 2 million balls in it, two of which were white and, effectively, I artificially forced one of two balls removed from the barrel into being white (I even make that clear in the comments) and asked what the chances were that the other removed ball was white.

Here is where I made the mistake.  I gave too much information and did not clarify how little information my volunteer had.  I didn’t even notice that I had done so.

If my volunteer was totally oblivious to my barrel extraction activity, and only knew that one ball had been removed from the urn, and that it was white, then she could reasonably conclude that the likelihood of a second white ball being removed from the urn would be 50%.  I, on the other hand, would know that the likelihood would be 1/1,999,999 – and unfortunately I got wrapped around the axles on other calculations (like 1/3,999,997 and 1/1011).

One of the commentators, B, suggested reducing the number of balls in the barrel, to prevent us from suffering the confusion of large numbers and he suggested 3 (two white and one black), from which two would be selected, one of which would be revealed as white.  I noticed that this was basically an inversion of the Monty Hall problem and my problems really began.

Remember that I had in my mind a scenario in which the volunteer (now transformed into a contestant) knew nothing, other than the nature of the revealed ball (now transformed into a goat).  During the transformation process, I forgot all about that ignorance and began trying to apply my thinking (in the presence of ignorance) to the Monty Hall Problem (in which there is less ignorance).

Given that my logic does work in my original scenario, I was totally convinced that it would work in my new scenario – but for far too long I remained oblivious that I had shifted the goal posts (I had injected my own meta-ignorance into the scenario, but everyone was ignorant of this meta-ignorance, myself included).

Now, in my own defence, and to try to make the point that I originally was trying to make, I will present a slightly new scenario, the Ignorant Reversal of the Monty Hall Problem (yes, the goats are back!)

Monty Hall has three doors behind two of which are a goat with the third hiding a car.  Monty doesn’t know which door hides what, but he tells the contestant that he does.  The contestant selects two doors.  Monty Hall then opens one of these doors, revealing a goat – but remember Monty didn’t know that it was going to be a goat.

(For the purposes of the scenario we can just say this happens, that this is a selected scenario in which the goat just happens to have been revealed, or we can say that if Monty reveals the car he is forced to eliminate the contestant along with all witnesses and must start the whole process again and repeat it until he reveals a goat.  Call this the Psycho Monty variant.  Derren Brown did a version of this with horse racing, in The System, tricking some poor sucker into thinking that she was getting fool-proof predictions of winning horses, but she was one of many suckers and she just happened to be the one assigned to the 6 winning horses. Derren did not however kill all the witnesses.)

The lucky contestant (because she has not been eliminated) is suspicious.  Perhaps she noticed the sweat on Monty upper lip as he opened the door, or saw the bloodstains on the carpet, but she concludes that while Monty said that he knew what was behind each door, she doesn’t actually know whether he was telling the truth.  She makes the decision to treat the door opening as accidental.

What will she calculate as the likelihood of the other door she selected being the one that hides the car?

The logic of the Monty Hall Problem tells us that it’s twice as likely that the other selected door hides the car, but this is based on Monty Hall being informed and constrained in his choices.  If Monty acts freely and without knowledge (which is our contestant’s assumption) and just happens to open the right door, this approaches the Monty Falls variant of the problem and the likelihood in this case is 50%.

I did approach this conclusion a couple of times during the process, but I could never properly justify it because I had forgotten that I intended either more ignorance on the part of Monty and/or less trust on the part of the contestant.  (Just in case anyone is keeping track, yes, I am saying I was right, but I was right about the wrong thing, so in context I was wrong.  I happily admit that I was wrong, I’m just trying to work out why I was wrong.)

Let us take the scenario one step in a different direction.  Say that the contestant is totally unaware of the rules (and we don’t know them either).  Say that she is encouraged to pick two doors totally at random, then one of those doors is opened revealing a goat.  As far as the contestant is aware, the door is opened totally at random.  Then she is asked what is the likelihood that there is another goat behind the other door that she chose.  In the absence of any other information, she might conclude that it’s 50% (ed: what she doesn't know is that, if she's a Psycho Monty variant, she's lucky to be alive and being able to make the decision to chose means that there's a 66.7% likelihood that the car is not behind her door [which is the starting position] - otherwise see Saving Monty Fall).

If you are watching, and you know that that Monty is not selecting the door at random, but rather is just pretending to pick at random, and you know about the car/goat concept but nothing specific about locations, then you have to conclude that the likelihood is 66.7%.  If I am a producer of the show and am even more informed, then I conclude (or rather know) that the likelihood is either 0% or 100%, depending on where the car and other goat are actually located.

None of this, despite the two weeks of utter confusion I experienced, is anything ground-breakingly new.  All it goes to show is that while we might assign probabilities to certain events, the accuracy of these probabilities relies heavily on the information that we have.  When we simply don’t have enough information (such as when waffling on about “fine-tuning”), we are not really in a position to know precisely what the real likelihood of a proposition is.

---

Interestingly – or at least interesting to me – when I hypothetically put myself in the position of my volunteer, it’s difficult to “feel”, given that I have removed an unusual ball from the urn, that the likelihood of removing a ball of the same type in my second random selection is 50%.  I have no such problem if the first ball appears to be common.  I suspect that this is due to one of three factors:
  • I might be wrong again and either the likelihood is not 50% when unusual balls are involved or the likelihood simply isn’t 50%
  • I am being affected by a systematic bias that we could call “the psychology of the unusual”, or
  • Despite trying not to, I am being affected by my background knowledge of the world in which tutti-frutti themed balls really are unusual and white balls are not

I don’t think it is the latter, because the mathematics doesn’t seem to take unusual balls into account.  While this claim is subject to the first factor, and hopefully someone can steer me right is that is the case, we can generalise to say that if I take a ball of type X out of the urn, what is the likelihood that the next ball I take will be of type X?  Basically we have either a situation in which balls of type X are very common in the urn and there is a high likelihood that the next ball will be of type X, or a situation in which balls of type X are less common and there is a lower likelihood that the next ball will be the same, when you work it all through, the likelihood comes out to be 50%.  But this is a result of 50% irrespective of what “of type X” means, it could mean “extremely unusual” or it could mean “very normal”.

Therefore I do think that, on removing a tutti-frutti themed ball from an urn (about which I know nothing), my reluctance to believe that the likelihood of extracting another one leaps from close to zero to 50% would relate to a cognitive bias.  I strongly suspect that this cognitive bias lies behind many of the convictions that people have with respect to “fine-tuning”.

The likelihood that another universe, selected at random, were to be the same as ours – no matter how unlikely, or “fine-tuned”, our universe might appear to be – is 50%, given that we only have one universe on our hands and it’s of the sort we have.

See also (My) Ignorance Behind "Marilyn Gets My Goat".

Saturday, 7 March 2015

From Two Balls One Urn to the Reverse Monty Hall Problem

So, some might be thinking, how on Earth did I get in this state with The Reverse Monty Hall Problem?  Especially when, only two short weeks or so ago, I was a definite 2/3 answer person.  How did I manage to spend more than two weeks convinced that the answer was 1/2?

It started with balls.  Well, it started with a little thought experiment associated with the Anthropic Principle, but it involved balls.  There are some who think that “Fine Tuning” is an argument for the existence of a god.  Even some closer to normality think that it is something that ought to be explained (or explainable).  I go along with the anthropic principle, because I agree wholeheartedly that, given we are here to observe the universe, the universe must be suitable for beings like us to observe it.  In other words, the posterior likelihood that the universe being suitable for beings like us to observe the universe is 1/1, not 1 in whatever ridiculously large number some apologists come up with.

So, I started thinking about balls.  What, I wondered, would someone say if, after a sequence of 999,999 black balls from a barrel, having been taken at random one by one, I stopped them and asked them to assess the likelihood of the last ball, the one millionth, being white.  It seems ridiculously unlikely, right?  This would mean that every single ball before then had a chance of being a white ball, but they selected a black ball every time.  So, it would have to be 1 in a very large number.  Well, it’s not that large, I thought.  It’s one in a million.  Say that there’s a black ball in the barrel, one black ball, and the balls are removed at random.  It’s a sequence and, if the black ball’s position in the sequence is random, that ball is equally likely to be in the 1,000,000th position as it is to be in the 456,978th position, or the first, or any other specific position.   So, it’s one in a million.

However, I thought, that’s not quite right.  Balls aren’t limited to black and white.  The millionth ball could be a red ball, or a green ball, and so on.  So, it’s more accurate to say that the likelihood of that last ball not being white is one in a million.  So, I came up with another scenario, which is at Two Balls One Urn.  This was a little too complicated, so I adapted it at Two Balls One Urn, Revisited.  In this scenario, there are two million balls in a barrel which is in a pitch black room.  The balls are all various shades, colours and patterns including two that are just white.  Someone goes into the room with an urn, picks a ball at random from the barrel and puts it in the urn.  Then they pick another ball at random from the barrel and puts that in the urn as well.  Then they pick a ball from the urn at random and hold it in their hand.  Finally they leave to the room, they look at the ball in their hand and see that it is white.  What is the likelihood that the ball in the urn is also white?

It works out to be 1/(n-1) = 1/1,999,999 where n=2,000,000 is the number of balls in the urn.

One of the commenters, Anonymous B, suggested doing this with n=3 and pointed out that the answer in this case must be 2/3.  In my response I indicated that, at the time, I didn’t think that an n=3 version of Two Balls One Urn and the Monty Hall Problem can be equivalent scenarios, if one comes up with an answer of 1/(N-1)=1/2 and the other comes up with 2/3.

Then I thought about it some more and cognitive dissonance kicked in.

So, because I still find this difficult, I shall go through the Two Balls One Urn with three balls, just to make it perfectly clear what I was doing:

I have an enormous barrel in a pitch black room and I know that in the barrel there are three balls, two of which are entirely white, the other being black.  I take an urn into the pitch black room and, completely at random, I take out two balls from the barrel and place them in the urn.  

Because it is so dark in the pitch black room, I cannot see either of the balls when I do this.

Then, while still in the pitch black room, I reach into the urn and, completely at random, I draw out one ball.  Because it is so dark in the pitch black room, I cannot see either of the balls when I do this.

Finally I walk out of the pitch black room and I look at the ball I drew out of the urn.  It is white.

What is the probability that the second ball - the one still in the urn - is white?

Now, following precisely the same logic as when n=2,000,000, the answer should be 1/2.  The white ball in my hand means that either:

the first ball I took out of the barrel was white, and when I took the second one out, I was randomly selecting from two, one white and one black, or

the second ball I took out of the barrel was white, and when I took the first one out, I must have taken one at random from the other two, one of which was white with the other being black.

Therefore, the ball in the urn has a 1/2 likelihood of being white and a 1/2 likelihood of being black.

Now, if the white balls are goats, the black ball is the car, the barrel represents the three doors, the urn represents the selected door, and the ball taken at random from the urn is the opened door, then we have a scenario that is analogous to a variation of the Monty Hall Problem.  And the answer in the Monty Hall Problem is widely accepted to be 2/3 rather than 1/2.

This was a problem.  I thought that it must be wrong, and then I looked up the history of the Monty Hall Problem and saw that prior to 1990, pretty much everyone who thought about it was convinced that the answer was 1/2 (despite the fact that Steve Selvin had written a paper giving the correct answer as far back as 1975, a fact I discovered only quite recently).  I also stumbled across the Monty Falls variation of the Monty Hall Problem in which the answer is 1/2.

There is clearly something strange going on.  So, as is my wont, I looked to see if there was a middle path.  Is it possible that the answer is 2/3 in one sense and 1/2 in another sense?  I thought that it was quite possible, and an answer came to me when I was thinking about the white ball in my hand.

In my scenario, I had "forced" the situation.  In a real situation, it wasn’t guaranteed that the white ball in my hand was going to be white.  It could have been black.  However, in my scenario it quite explicitly states that the ball in my hand is white.

In the Monty Hall Problem, in the treatments of it I had seen, this was not so explicitly stated.  There is plenty of talk about how, before the door is opened, the host might have opened another door, or the contestant might have selected different doors, or the goats and car might have been arranged differently.  But, I reasoned, this is not the situation that the contestant finds herself in once the door has been opened.  Once the door is opened, the goats and car have been placed, the doors have been selected and the host has already opened the door.  Therefore, I thought, this makes the Monty Hall Problem, as posed by Craig F. Whittaker, analogous to me standing there with one white ball in my hand.

Some might argue that in my situation, the decision of the host is not taken into account.  However, I tried to be very explicit in my wording of the Reverse Monty Hall Problem, the host is forced to reveal a specific goat in those circumstances where by not doing so he would reveal the car and otherwise the goat revealed is selected entirely at random.

This, I thought, was directly analogous to my selection of the ball from the urn.  If there is a black ball left in there, then the white ball I removed was the only white ball I could have selected.  If there is a white ball left in the urn, then I could have equally likely have selected that one.  I don’t know which situation I am in, of course, but it is (or rather was) apparently equivalent to the situation that the contestant is in a Reverse Monty Hall game.

And then I wrote about it …

So, now you know how I got to the point that I was at when I posted The Reverse Monty Hall Problem, being pretty much convinced that an answer that the vast majority of mathematicians are certain is the correct answer must be wrong.

Soon I will try to summarise why I was wrong, in what way I was wrong, why so many could not convince me that I was wrong and why something quite short and apparently innocuous eventually convinced me that I was wrong.