Tuesday, 27 March 2012

Judea Pearl wins Turing Award

Judea Pearl, who has done more than anybody to develop Bayesian networks and causal reasoning, has won the 2011 Turing Award for work on AI reasoning.

We are also delighted to announce that Judea has written the Foreword for our forthcoming book.

The danger of p-values and statistical significance testing

I have just come across an article in the Financial Times (it is not new - it was published in 2007) titled "The Ten Things Everyone Should Know About Science".  Although the article is not new the source where I found the link to it is, namely right at the top of the home page for the 2011-12 course on Probabilistic Systems Analysis at MIT. In fact the top bullet point says:
The concept of statistical significance (to be touched upon at the end of this course) is considered by the Financial Times as one of " The Ten Things Everyone Should Know About Science".
The FT article does indeed list "Statistical significance" as one of the ten things, along with: Evolution, Genes and DNA, Big Bang, Quantum Mechanics, Relativity, Radiation, Atomic and Nuclear Reactions, Molecules and Chemical Reactions, and Digital data.   That is quite illustrious company, and in the sense that it helps promote the importance of correct probabilistic reasoning I am delighted. However, as is fairly common, the article assumes that 'statistical sugnificance' is synonymous with p-values. The article does hint at the fact that there there might be some scientists who are sceptical of this approach when it says:
Some critics claim that contemporary science places statistical significance on a pedestal that it does not deserve. But no one has come up with an alternative way of assessing experimental outcomes that is as simple or as generally applicable.
In fact, that first sentence is a gross under-statement, while the second is simply not true. To see why the first sentence is a gross understatement look at this summary (which explains what p-values are) that appears in Chapter 1 of our forthcoming book (you can see full draft chapters of the book here). To see why the second sentence is not true look at this example from Chapter 5 of the book (which also shows why Bayes offers a much better alternative). Also look at this (taken from Chapter 10) which explains why the related 'confidence intervals' are not what most people think (and how this dreadful approach can also be avoided using Bayes).

Hence it is very disappointing that an institute like MIT should be perpetuating the myths about this kind of significance testing. The ramifications of this myth have had (and continues to have) a profound negative impact on all empirical research. The book "The Cult of Statistical Significance: How the Standard Error Costs Us Jobs, Justice, and Lives (Economics, Cognition & Society)" by Ziliak and McCloskey (The University of Michigan Press, 2008) provides extensive evidence of flawed studies and results published in reputable journals across all disciplines. It is also worth looking at the article "Why Most Published Research Findings Are False". Not only does it mean that 'false' findings are published but also that more scientifically rigorous empirical studies are rejected because authors have not performed the dreaded significance tests demanded by journal editors or reviewers.  This is something we see all the time and I can share an interesting anecdote on this. I was recently discussing a published paper with its author. The paper was specifically about using the Bayesian Information Criteria to determine which model was producing the best prediction in a particular application. The Bayesian analysis was the 'significance test' (only a lot more informative).Yet at the end of the paper was a section with a p-value significance test analysis that was redundant and uninformative. I asked the author why she had included this section as it kind of undermined the value of the rest of the paper. She told me that the paper she submitted did not have this section but that the journal editors had demanded a p-value analysis as a requirement for publishing the paper.

Thursday, 26 January 2012

Prosecutor Fallacy in Stephen Lawrence case?

If you believe the 'respectable' media reporting of the Stephen Lawrence case then there were at least two blatant instances of the prosecutor fallacy committed by expert forensic witnesses.

For example the BBC reported forensic scientist Edward Jarman as saying that
"..the blood stain on the accused's jacket was caused by fresh blood with a one-in-a-billion chance of not being victim Stephen Lawrence's."
Numerous other reports contained headlines with similar claims about this blood match. The Daily Mail (above) states explicitly:
".. the chances of that blood belonging to anyone but Stephen were rated at a billion to one, the Old Bailey heard."
Similar claims were made about the fibres found on the defendants' clothing that matched those from Stephen Lawrence's clothes.

If the expert witnesses had indeed made the assertions as claimed in the media reports then they would have committed a well known probability fallacy which, in this case, would have grossly exaggerated the strength of the prosecution evidence. The fallacy - called the prosecutor fallacy - has a long history (see, for example our reports here and here as well as summary explanation here).  Judges and lawyers are expected to ensure the fallacy is avoided because of its potential to mislead the jury. In practice the fallacy continues to be made in courts. However, despite the media reports the fallacy was not made in the Stephen Lawrence trial. The experts did not make the assertions that the media claimed they made. Rather, it was the media who misunderstood what the experts said and it was they who made the prosecutor fallacy. What is still disturbing is that, if media reporters with considerable legal knowledge made the fallacy then it is almost certain that the jury misinterpreted the evidence in the same way. In other words, although the prosecutor fallacy was not stated in court, it may still have been made in the jury's decision-making.

To understand the fallacy and its impact formally, suppose:
  • H is the hypothesis "the blood on Dobson's jacket did not come from Stephen Lawrence"
  • E is the evidence "blood found matches Lawrence's DNA".

What the media stated is that the forensic evidence led to the probability of H (given E) being a billion to one.

But, in fact, the forensic experts did not (and could not) conclude anything about the probability of H (given E). What the media have done is confuse this probability with the (very different) probability of E given H.

What the experts were stating was that (providing there was no cross contamination or errors made) the probability of E given H is one in a billion. In other words what the experts were asserting was

"The blood found on the jacket matches that of Lawrence and such a match is found in only one in a billion people. Hence the chances of seeing this blood match evidence is one in a billion if the blood did not come from Lawrence".

In theory (since there are about 7 billion people on the planet) there should be about 7 people (including Lawrence) who would have the same matching blood DNA to Lawrence. If none of the others could be ruled out as the source of the blood on the jacket then the probability of H given E is not one in a billion as stated by the media but 6 out of 7. This highlights the potential enormity of the fallacy.

Even if we could rule out everyone who never came into contact with Dobson that would still leave, say, 1000 people. In that case the probability of H given E is about one in a million. That is, of course, a very small probability but the point is that it is a very different probability to the one the media stated.

The main reason why the fallacy keeps on being repeated (by the media at least) in these kind of cases is that people cannot see any real difference between a one in a billion probability and a one in a million probability (even though the latter is 1000 times more likely). They are both considered 'too small'.

Finally, it is also important to note that the probabilities stated were almost meaningless because of the simplistic assumptions (grossly favourable to the prosecution case) that there was no possibility of either cross-contamination of the evidence or errors in its handling and DNA testing. The massive impact such error possibilities have on the resulting probabilities is explained in detail here.

Tuesday, 8 November 2011

Nonsensical probabilities about asteroid risk

There is an article in today's Evening Standard in which, rather depressingly, someone who should know better (Roger Highfield, Editor of New Scientist) reels off a typically misleading probability about asteroid strike risk (this is in response to the news today that an asteroid was within just 200,000 miles of Earth).

Quoting a recent book by Ted Nield he says (presumably to comfort readers) that

Our chances of dying as a result of being hit by a space rock are something like one in 600,000. 
There are all kinds of  ambiguities about the statement that I won't go into (involving assumptions about  random people of  'average' age and 'average' life expectancy) but even ignoring all that, if the statement is supposed to be a great comfort to us then it fails miserably. That is because it can reasonably be interpreted as providing the 'average' probability that a random person living today will eventually die from being hit by a space rock. Assuming a world population of 7 billion that's about 12,000 of us. And 12,000 actual living people is a pretty large number to die in this way. But it is about as meaningful as putting Arnold Schwarzenegger in a line up with a thousand ants and asserting that the average height of those in the line-up is 3 inches tall. The key issue here is that large asteroid strikes are, like Schwarzenegger in the line-up, low probability high impact events. Space rocks will not kill a few hundred people every year as implied by the original statement, just as there are no 3-inch tall beings in the line-up. Tens of thousands of years pass between them killing any more than a handful of people. But eventually one will wipe out most of the world's population. 

What Ted Nield should have stated (and what we are most interested in knowing) was the probability that a large space rock (one big enough to cause massive loss of life) will strike Earth in the next 50 years.

Indeed, I suspect that (using Nield's own analysis) this probability would be close to the 1 in 600,000 quoted (given that incidents of small space rocks killing a small number of people are very rare). You might argue I am splitting hairs here but there is an important point of principle. Nield and Highfield avoid stating an explicit probability of a very rare event (such as in this case a massive asteroid strike) because there is a natural resistance (especially from non-Bayesians) to do so. For some reason it is more palatable in their eyes to consider the probability of a random person dying (albeit due to a rare event), presumably because it can more easily be imagined. But, as I have hopefully shown, that only creates more confusion.

Friday, 4 November 2011

Bayes and the Law: Nature article

My commentary piece on the role of Bayes in the Law has just appeared in Nature. A pdf of an extended draft on which it was based is here.

Monday, 3 October 2011

Bayes and the Law: Guardian article

There is an article in the Guardian today that is the result of an interview I had with the journalist Angela Saini. It is about the issue of Bayes and the Law following the RvT ruling and it includes a reference to the consortium that we are putting together to improve the situation.

Friday, 2 September 2011

Another specialist risk assessment company gets bought out

Algorithmics, a company specialising in risk software for financial institutions, has been sold to IBM for $387m.

Agena partnered Algorithmics during the period 2003-2005 when there was a clamour for so-called 'advanced measurement approaches' to operational risk assessment. The Basel 2 accord specified that banks which used a validated advanced measurement approach to calculate their operational risk exposure could set aside a lower percentage capital allocation. This meant there was a major financial incentive for banks to develop such approaches. Most banks looked at Bayesian networks as a potential solution and, indeed, a number wanted Algorithmics to provide such a solution to integrate with the existing credit and market risk software that Algo provided. Algorithmics had started work on their own Bayesian network platform for OpRisk, but decided that AgenaRisk was superior. Hence we partnered them in projects with some major banks, with Agena providing the underlying Bayesian network technology and modelling skills and Algorithmics providing the reporting infrastructure.  Here is a paper we wrote that gives a feel for the BN approach to operational risk that we developed.

In late 2005 Algorithimcs actually got taken over by the Fitch Group. Since Fitch already had their own (non-Bayesian) OpRisk solution - which they had massively invested in - the partnership effectively ended then, as it appears did Algorithmics' interest in Bayesian networks. This is a great shame, especially when you consider the mess that financial institutions have made using classical statistics.

It is difficult to determine the extent to which banks are using Bayesian networks but, as we described here, there are plenty of financial analysts who are using fundamentally flawed methods in situations when the Bayesian approach would work.