Sunday, 1 January 2017

The problem with the likelihood ratio for DNA mixture profiles


We have written many times before (see the links below) about use of the Likelihood Ratio (LR) in legal and forensic analysis.

To recap: the LR is a very good and simple method for determining the extent to which some evidence (such as DNA found at the crime scene matching the defendant) supports one hypothesis (such as "defendant is the source of the DNA") over an alternative hypothesis (such as "defendant is not the source of the DNA"). The previous articles discussed the various problems and misinterpretations surrounding the use of the LR. Many of these arise when the hypotheses are not mutually exclusive and exhaustive. This problem is especially pertinent in the case of 'DNA mixture' evidence, i.e. when some DNA sample relevant to a case comes from more than one person. With modern DNA testing techniques it is common to find DNA samples with multiple (but unknown number of) contributors. In such cases there is no obvious 'pair' of hypotheses that are mutually exclusive and exhaustive, since we have individual hypotheses such as:
  • H1: suspect + one unknown
  • H2: suspect + one known other 
  • H3: two unknowns
  • H4: suspect + two unknowns 
  • H5: suspect + one known other + one unknown
  • H6: suspect + two known others
  • H7: three unknowns 
  • H8: one known other + two unknowns
  • H9: two known others + one unknown
  • H10: three known others
  • H11:  suspect + three unknowns 
  • etc.
It is typical in such situations to focus on the 'most likely' number of contributors (say n) and then compare the hypothesis "suspect + (n-1) unknowns" with the hypothesis "n unknowns". For example, if there are likely to be 3 contributors then typically the following hypotheses are compared:
  • H1: suspect + two unknowns
  • H2: three unknowns
Now, to compute the LR we have to compute the likelihood of the particular DNA trace evidence E under each of the hypotheses. Generally both of these are extremely small numbers, i.e. both the probability values P(E | H1) and P( E | H2) are very small numbers. For example, we might get something like
  • P(E | H1) = 0.00000000000000000001  (10 to the minus 20)
  • P(E | H2) = 0.00000000000000000000000001  (10 to the minus 26)
For a statistician, the size of these numbers does not matter – we are only interested in the ratio (that is precisely what the LR is) and in the above example the LR is very large (one million) meaning that the evidence is a million times more likely to have been observed if H1 is true compared to H2. This seems to be overwhelming evidence that the suspect was a contributor. Case closed?

Apart from the communication problem in court of getting across what this all means (defence lawyers can and do exploit the very low probability of E given H1) and how it is computed, there is an underlying statistical problem with small likelihoods for non-exhaustive hypotheses and I will highlight the problem with two scenarios involving a simple urn example. Superficially, the scenarios seem identical. The first scenario causes no problem but the second one does. The concern is that it is not at all obvious that the DNA mixture problem always corresponds more closely to the first scenario than the second.

In both scenarios we assume the following:

There is an urn with 1000 balls – some of which are white. Suppose W is the (unknown) number of white balls. We have 2 hypotheses:
  • H1: W=100
  • H2:  W=90
We can draw a ball as many times as we like, note its colour and replace it (i.e. sample with replacement). We wish to use the evidence of 10,000 such samples.

Scenario 1: We draw 1001 white balls. In this case using standard statistical assumptions we calculate P(E | H1) = 0.013, P(E|H2) = 0.0000036. Both values are small but the LR is large, 3611, strongly favouring H1 over H2.

Scenario 2: We draw 1100 white balls. In this case P(E | H1) = 0.000057, P(E|H2) < 0.00000001. Again both values are very small but the LR is very large, strongly favouring of H1 over H2.

(note: in both cases we could have chosen a much larger sample and got truly tiny likelihoods but these values are sufficient to make the point).

So in what sense are these two scenarios fundamentally different and why is there a problem?

In scenario 1 not only does the conclusion favouring H1 make sense, but the actual number of balls drawn is very close to the expected number we would get if H1 were true (in fact, W=100 is the 'maximum likelihood estimate' for number of balls). So not only does the evidence point to H1 over H2, but also to H1 over any other hypothesis (and there are 1000 different hypotheses W=0, W=1, W=2 etc.).

In scenario 2 the evidence is actually even much more supportive of H1 over H2 than in scenario 1. But it is essentially meaningless because it is virtually certain that BOTH hypotheses are false.

So, returning to the DNA mixture example, it is certainly not sufficient to compare just two hypotheses. The LR of one million in favour of H1 over H2 may be hiding the fact that neither of these hypotheses is true. It is far better to identify as exhaustive a set of hypotheses as is realistically possible and then determine the individual likelihood value of each hypothesis. We can then identify the hypothesis with the highest likelihood value and consider its LR compared to each of the other hypotheses.

Tuesday, 8 November 2016

Confusion over the Likelihood Ratio


The 'Likelihood Ratio' (LR) has been dominating discussions at the third workshop  in our Isaac Newton Institute Cambridge Programme Probability and Statistics in Forensic Science.
There have been many fine talks on the subject - and these talks will be available here for those not fortunate enough to be attending.

We have written before (see links at bottom) about some concerns with the use of the LR. For example, we feel there is often a desire to produce a single LR even when there are multiple different unknown hypotheses and dependent pieces of evidence (in such cases we feel the problem needs to be modelled as a Bayesian network)- see [1]. Based on the extensive discussions this week, I think it is worth recapping on another one of these concerns (namely when hypotheses are non-exhaustive).

To recap: The LR  is a formula/method that is recommended for use by forensic scientists when presenting evidence - such as the fact that DNA collected at a crime scene is found to have a profile that matches the DNA profile of a defendant in a case. In general, the LR can a very good and simple method for communicating the impact of evidence (in this case on the hypothesis that the defendant is the source of the DNA found at the crime scene).

To compute the LR, the forensic expert is forced to consider the probability of finding the evidence under both the prosecution and defence hypotheses. So, if the prosecution hypothesis Hp is "Defendant is the source of the DNA found" and the defence hypothesis Hp is "Defendant is not the source of the DNA found" then we compute both the probability of the evidence given Hp - written P(E | Hp) - and the probability of the evidence given Hd - written P(E | Hd). The LR is simply the ratio of these two likelihoods, i.e. P(E | Hp) divided by P(E | Hd).

The very act of considering both likelihood values is a good thing to do because it helps to avoid common errors of communication that can mislead lawyers and juries (notably the prosecutor's fallacy). But, most importantly, the LR is a measure of the probative value of the evidence. However, this notion of probative value is where misunderstandings and confusion sometimes arise. In the case where the defence hypothesis is the negation of the prosecution hypothesis (i.e. Hd is the same as "not Hp" as in our example above) things are clear and very powerful because, by Bayes theorem:
  • when the LR is greater than one the evidence E supports the prosecution hypothesis in the real sense that the posterior probability of Hp increases, i.e. P(Hp | E) > P(Hp)  and, hence the posterior probability of Hd decreases by the same amount. In fact the posterior odds of the prosecution hypothesis increase by a factor of LR over the prior odds.
  • when the LR is less than one it supports the evidence E supports the defence hypothesis in the real sense that the posterior probability of Hd increases, i.e. P(Hd | E) > P(Hd)  and, hence the posterior probability of Hp decreases by the same amount. In fact  the posterior odds of the defence hypothesis increase by a factor of LR over the prior odds.
  • when the LR is equal to one then the evidence supports neither hypothesis and so is 'neutral' in the real sense that the posterior probability of both H and Hd are unchanged,  i.e. P(Hp | E) = P(Hp) and P(Hd | E) = P(Hd). In such cases, since the evidence has no probative value lawyers and forensic experts believe it should not be admissible.
Hence, in this case, providing that we know what the prior probability of Hp is we are able to compute the posterior probability of Hp.

However, things are by no means as clear and powerful when the hypotheses are not exhaustive (i.e. the negation of each other) and in most forensic applications this is the case. For example, in the case of DNA evidence, while the prosecution hypothesis Hp is still "defendant is source of the DNA found" in practice the defence hypothesis Hd is often something like "a person unrelated to the defendant is the source of the DNA found".

In such circumstances the LR can only help us to distinguish between which of the two hypotheses is better supported by the evidence. Unlike the case for exhaustive hypotheses a LR greater than 1 does not necessarily mean that the evidence 'supports' the prosecution hypothesis. In fact, the LR can be very large - i.e. the evidence strongly supports the prosecution hypothesis over the defence hypothesis - even though the posterior probability of the prosecution hypothesis goes down.  This rather worrying point is not understood by all forensic scientists (or indeed by all statisticians). Consider the following example (it's a made-up coin tossing example, but has the advantage that the numbers are indisputable):
Fred claims to be able to toss a fair coin in such a way that about 90% of the time it comes up Heads. So the main hypothesis is
  Hp: Fred has genuine skill
To test the hypothesis, we observe him toss a coin 10 times. It comes out Heads each time. So our evidence E is 10 out of 10 Heads. Our alternative hypothesis is:
  Hd: Fred is just lucky.

By Binomial theorem assumptions, P(E | Hp) is about 0.35 while P(E | Hd) is about 0.001. So the LR is about 350 - that is strongly in favour of Hp.

However, the problem here is that Hp and Hd are not exhaustive. There could be another hypotheses Hx: "Fred is cheating by using a double-headed coin". Now, P(E | Hx) = 1.

If we assume that Hp, Hd and Hx are the only possible hypotheses* (i.e. they are exhaustive) and that the priors are equally likely, i.e. each is equal to 1/3 then the posteriors after observing the evidence E are:

Hp: 0.25907      Hd: 0.00074          Hx: 0.74019

So, after observing the evidence E, the posterior for Hp has actually decreased despite the very large LR in its favour over Hd.
In the above example, a good forensic scientist - if considering only Hp and Hd - would conclude by saying something like
"The evidence is 350 times more likely under hypothesis Hp than Hd, but tells us nothing about whether we should have greater belief in Hp being true; indeed, it is possible that the evidence may more strongly support some other hypothesis not considered and even make our belief in Hp decrease". 
However, in practice (and I can confirm this from having read numerous DNA reports) no such careful statement is made. In fact, the most common assertion used in such circumstances is:
 "The evidence provides strong support for hypothesis Hp"
Such an assertion is not only mathematically wrong but highly misleading. Consider, as discussed above, a DNA case where:

 Hp is "defendant is source of the DNA found"
 Hd is  "a person unrelated to the defendant is the source of the DNA found".

This particular Hd hypothesis is a common convenient choice for the simple reason that P(E | Hd) is relatively easy to compute (it is the 'random match probability'). For single-source, high quality DNA this probability can be extremely small - of the order of one over several billions; since P(E | Hp) is equal to 1 in this case the LR is several billions. But, this does NOT provide overwhelming support for Hp as is often assumed unless we have been able to rule out all relatives of the defendant as suspects. Indeed, for less than perfect DNA samples it is quite possible for the LR  to be in the order of millions but for a close relative to be a more likely source than the defendant.

While confusion and misunderstandings can and do occur as a result of using hypotheses that are not exhaustive, there are many real examples where the choice of such non-exhaustive hypotheses is actually negligent.  The following appalling example is based on a real case (location details changed as an appeal is ongoing):
The suspect is accused of committing a crime in a particular rural location A near his home village in Dorset. The evidence E is soil found on the suspect's car.  The prosecution hypothesis Hp is "the soil comes from A". The suspect lives (and drives) near this location but claims he did not drive to that specific spot. To 'test' the prosecution hypothesis a soil expert compares Hp with the hypothesis Hd: "the soil comes from a different rural location". However, the 'different rural location'  B happens to be 500 miles away in Perth Scotland (simply because it is close to where the soil analyst works and he assumes soil from there is 'typical' of rural soil). To carry out the test the expert considers soil profiles of E and samples from the two sites A and B.

Inevitably the LR strongly favours Hp (i.e. site A)  over Hd (i.e. site B); the soil profile on the car - even if it was never at location A - is going to be much closer to the A profile than the B profile. But we can conclude absolutely nothing about the posterior probability of A. The LR is completely useless - it tells us nothing other than the fact that the car was more likely to have been driven in the rural location in Dorset than in a a rural location in Perth. Since the suspect had never driven the car outside Dorset this is hardly a surprise.  Yet, in the case this soil evidence was considered important since it was wrongly assumed to mean that it "provided support for the prosecution hypothesis".
This example also illustrates, however, why in practice it can be impossible to consider exhaustive hypotheses. For such soil cases, it would require us to consider samples from every possible 'other' location. What an expert like Pat Wiltshire (who is also a participant on the FOS programme) does is to choose alternative sites close to the alleged crime scene and compare the profile of each of those and the crime scene profile with the profile from the suspect. While this does not tell us if the suspect was at the crime scene it can tell us how much more likely the suspect was to have been there rather than sites nearby.

*as pointed out by Joe Gastwirth there could be other hypotheses like "Fred uses the double-headed coin but switches to a regular coin after every 9 tosses"

References
  1. Fenton N.E, Neil M, Berger D, “Bayes and the Law”, Annual Review of Statistics and Its Application, Volume 3, 2016 (June), pp 51-77 http://dx.doi.org/10.1146/annurev-statistics-041715-033428 .Pre-publication version here and here is the Supplementary Material See also blog posting.
  2. Fenton, N. E., D. Berger, D. Lagnado, M. Neil and A. Hsu, (2013). "When ‘neutral’ evidence still has probative value (with implications from the Barry George Case)", Science and Justice, http://dx.doi.org/10.1016/j.scijus.2013.07.002.  A pre-publication version of the article can be found here.

See also previous blog postings:



Friday, 7 October 2016

Bayesian Networks and Argumentation in Evidence Analysis


Some of the workshop participants
On 26-29 September 2016 a workshop on "Bayesian Networks and Argumentation in Evidence Analysis" took place at the Isaac Newton Institute Cambridge. This workshop, which was part of the FOS Programme was also the first public workshop of the ERC-funded project Bayes-Knowledge (ERC-2013-AdG339182-BAYES_KNOWLEDGE).

The workshop was a tremendous success, attracting many of the world's leading scholars in the use of Bayesian networks in law and forensics. Most of the presentations were filmed and can now be viewed here.

There was also a pre-workshop meeting on 23-24 September where participants focused on an important Dutch case that recently went to appeal. The partcipants were divided into two groups - one group developed a BN model of the case and the other developed an agumentation/scenarios-based model of the case. We plan to further develop these and write up the results.

Some of the participants at the pre-workshop meeting anyalysing a specific Dutch case


The Bayesian Networks mutual exclusivity problem

Several years ago when we started serious modelling of legal arguments using Bayesian networks we hit a problem that we felt would be easily solved. We had a set of mutually exclusive events such as "X murdered Y, Z murdered Y, Y was not murdered" that we needed to model as separate variables because they had separate causal pathways and evidence.

It turned  out that existing BN modelling techniques cannot capture the correct intuitive reasoning when a set of mutually exclusive events need to be modelled as separate nodes instead of states of a single node. The standard proposed ’solution’, which introduces a simple constraint node that enforces mutual exclusivity, fails to preserve the prior probabilities of the events and is therefore flawed.

In 2012 myself (and the co-authors listed below) produced an initial novel and simple solution to this problem that works in a reasonable set of circumstances, but it proved to be difficult to get people to understand why the problem was an important one that needed to be solved. After many changes and iterations this work has finally been published and, as a 'gold access paper' it is free for anybody to download in full (see link below).

During the current Programme "Probability and Statistics in Forensic Science" that I am helping to run at the Isaac Newton Institute for Mathematical Sciences, Cambridge, 18 July - 21 Dec 2016, it has become clear that the mutual exclusivity problem is critical in any legal case where there are diverse prosecution and defence narratives. Although our solution does not work in all cases (and indeed we are working on more comprehsive approaches) we feel it is an important start.

Norman Fenton, Martin Neil, David Lagnado, William Marsh, Barbaros Yet, Anthony Constantinou, "How to model mutually exclusive events based on independent causal pathways in Bayesian network models", Knowledge-Based Systems, Available online 17 September 2016
http://dx.doi.org/10.1016/j.knosys.2016.09.012

Saturday, 17 September 2016

Bayesian networks: increasingly important in cross disclipinary work

The growing importance of Bayesian networks was demonstrated this week by the award of a prestigious Leverhulme Trust Research Project Grant of £385,510 to Queen Mary University of London that ultimately will lead to improved design and use of self-monitoring systems such as blood sugar monitors, home energy smart meters, and self-improvement mobile phone apps.

The project, CAUSAL-DYNAMICS ("Improved Understanding of Causal Models in Dynamic Decision-making") is a collaborative project, led by Professor Norman Fenton of the School of Electronic Engineering and Computer Science, with co-investigators Dr Magda Osman (School of Biological and Chemical Sciences), Prof Martin Neil (School of Electronic Engineering and Computer Science) and Prof David Lagnado (Department of Experimental Psychology, University College London).

The project exploits Fenton and Neil's expertise in causal modelling using Bayesian networks and Osman and Lagnado's expertise in cognitive decision making. Previously, psychologists have extensively studied dynamic decision-making without formally modelling causality while statisticians, computer scientists, and AI researchers have extensively studied causality without considering its central role in human dynamic decision making. This new project starts with the hypothesis that we can formally model dynamic decision-making from a causal perspective. This enables us to identify both where sub-optimal decisions are made and to recommend what the optimal decision is. The hypothesis will be tested in real world examples of how people make decisions when interacting with dynamic self-monitoring systems such as blood sugar monitors and energy smart meters and will lead to improved understanding and design of such systems.

The project is for 3 years starting Jan 2017. For further details, see: CAUSAL-DYNAMICS.

WATCH THIS SPACE FOR THE ANNOUNCEMENT VERY SOON OF TWO OTHER MAJOR NEW CROSS-DISCIPLINARY BAYESIAN NETWORK PROJECTS!!

About the Leverhulme Trust
The Leverhulme Trust was established by the Will of William Hesketh Lever, the founder of Lever Brothers. Since 1925 the Trust has provided grants and scholarships for research and education; today it is one of the largest all-subject providers of research funding in the UK, distributing approximately £80 million a year. For more information: www.leverhulme.ac.uk / @LeverhulmeTrust

Friday, 16 September 2016

Bayes and the Law: what's been happening in Cambridge and how you can see it


Programme Organisers (left to right): Richard Gill, David Lagnado, Leila Schneps, David Balding, Norman Fenton
Since 21 July 2016 I have been running the Isaac Newton Institute (INI) Programme on Probability and Statistics in Forensic Science in Cambridge.

For those of you who were not fortunate enough to be at the first formal workshop "The nature of questions arising in court that can be addressed via probability and statistical methods" (30 August to 2 September) you can watch the full videos here of most of the 35 presentations on the INI website. The presentation slide are also available in the INI link..

The workshop attracted many of the world's leading figures from the law, statistics and forensics with a mixture of academics (including mathematicians and legal scholar), forensic practitioners, and practicing lawyers (including judges and eminent QCs). It was rated a great success.

The second formal workshop "Bayesian Networks and Argumentation in Evidence Analysis" will take place on 26-29 September. It is also part of the BAYES-KNOWLEDGE project programe of work. For those who wish to attend, but cannot, the workshop will be streamed live.

Norman Fenton, 16 September 2016

Links



Friday, 1 July 2016

The likelihood ratio and why its use in forensic analysis is often flawed

FORREST 2016 (for details see here)

I am giving the opening address at the Forensic Institute 2016 Conference (FORREST 2016) in Glasgow on 5 July 2016. The talk is about the benefits and pitfalls of using the likelihood ratio to help understand the impact of forensic evidence. The powerpoint slide show for my talk is here.

While a lot of the material is based on our recent Bayes and the Law paper, there is a new simple example of the danger of using the likelihood ratio (LR) when the defence hypothesis is not the negation of the prosecution hypothesis. Recall that the LR for some evidence E is the probability of E given the prosecution hypothesis divided by the probability of E given the defence hypothesis. The reason the LR is popular is because it is a measure of the probative value of the evidence E in the sense that:
  • LR>1 means E supports the prosecution hypothesis
  • LR<1 means  E supports the defence hypothesis
  • LR=1 means E has no probative value
This follows from Bayes Theorem but only when the defence hypothesis is the negation of the prosecution hypothesis. The problem is that there are Forensic Science Guidelines* that explicitly state that this requirement is not necessary. But if the requirement is not met then it is possible to have LR<1 even though E actually supports the prosecution hypothesis. Here is the example:


A raffle has 100 tickets numbered 1 to 100

Joe buys 2 tickets and gets numbers 3 and 99

The ticket is drawn but is blown away in the wind.

Joe says he had the winning ticket, but the organisers say neither 3 nor 99 was not the winning ticket. In this case the prosecution hypothesis H is “Joe won the raffle” (i.e. either 3 or 99 was the winning ticket).
Suppose we have the following evidence E presented by a totally reliable eye witness:
 
E: “winning ticket was an odd nineties number (i.e. 91, 93, 95, 97, or 99)”

Does the evidence E support H? let's do the calculations:
  • Probability of E given H = 1/2
  • Probability of E given not H = 4/98
So the LR  is  (1/2)/(4/98) = 12.25

That means the evidence CLEARLY supports H. In fact, the probability of H increases from a prior of 1/50 to a posterior of 1/5, so there is no doubt it is supportive.

But suppose the organisers’ assert that their (defence) hypothesis is:

H’: “Winning ticket was a number between 95 and 97”

Then in this case we have:
  • Probability of E given H = ½
  • Probability of E given H’ = 2/3
So the LR  is  ( 1/2)/(2/3) = 0.75

That means that in this case the evidence supports H’ over H. The problem is that, while the LR does indeed 'prove' that the evidence is more supportive of H' than H that is actually irrelevant unless there is other evidence that proves that H' is the only possible alternative to H (i.e. that H' equivalent to 'not H').  In fact, the  'defence' hypothesis has been cherry picked. The evidence E supports H irrespective of which cherry-picked alternative is considered.  
Norman Fenton, 1 July 2016
 

*Jackson G, Aitken C, Roberts P. 2013. Practitioner guide no. 4. Case assessment and interpretation of expert evidence: guidance for judges, lawyers, forensic scientists and expert witnesses. London: R. Stat. Soc. http://www.maths.ed.ac.uk/∼cgga/Guide-4-WEB.pdf.  Page 29: "The LR is the ratio of two probabilities, conditioned on mutually exclusive (but not necessarily exhaustive) propositions."

See also: