Thursday, 24 September 2020

Don’t Panic: Limits to what we know about UK Covid-19 PCR testing, inferred infection rates and the rate of false positives

A short paper by Neil et al uses Bayesian analysis to examine the latest (up to 22 Sept) Covid data to determine whether there is evidence to support the Government claim of an exponential 'second wave'. Concludes there remains insufficiently solid evidence, despite the number of tests done, to support any claim there is an exponential increase. There is no reason to panic.

Infection prevalence between April and September (Week 1 is 12 Aug and week 39 is 19 Sept

Link to paper:

https://qm-rim.org/wp-content/uploads/2020/09/Neil-et-al-2020-Limits-to-UK-PCR-testing.pdf 

 

See also:

Plotting new Covid cases per 1000 tested

Monday, 21 September 2020

New paper highlights serious limitations of using the likelihood ratio for mixture DNA profile matching

I have written many times before on this blog about the benefits and (extreme) limitations of using the likelihood ratio (LR) to determine the strength of forensic match evidence - such as DNA evidence. The LR is the probability of the evidence (i.e. 'finding a match') under the prosecution hypothesis divided by the probability of the evidence under the defence hypothesis.

When a forensic expert calculates a high LR for the prosecution hypothesis "DNA found at the crime scene matches the DNA profile of the suspect" they typically conclude that ‘this provides strong support for the prosecution hypothesis’. However, it is well known that a high LR does not necessarily mean a high (posterior) probability that the hypothesis is true because this depends on the prior probability of the hypothesis. Assuming that it does is an example of what is called the prosecutor's fallacy; judges are expected to warn both lawyers and expert witnesses when this mistake is made in court. What is less well understood is that, in order to draw any rational conclusions at all about the probative value of a high LR, the defence hypothesis used to determine the LR has to be the negation of the prosecution hypothesis (formally we say the hypotheses must be mutually exclusive and exhaustive). So, in the above example the defence hypothesis has to be ‘DNA does not come from the suspect’. Yet in practice, it is common to use ‘DNA comes from a person unrelated to the defendant’ which of course is not the negation of the prosecution hypothesis. When the defence hypothesis is not the negation, then a high LR could actually mean the exact opposite of what is claimed: it could provide more support for the hypothesis that the DNA does not come from the suspect than it does for the prosecution hypothesis.

While the problem of using the LR with hypotheses that are not mutually exclusive and exhaustive is known (but not widely understood) - for 'single' DNA profiles (i.e. those where the DNA found can only possibly come from a single person) it is even more serious for DNA mixture profiles (i.e those where the DNA is a mixture of at least two people). However, for mixture profiles, not having mutually exclusive and exhaustive hypotheses when using the LR is just one of several serious problems. A new paper shows the extent to which a very high LR for a ‘match’ from a DNA mixture profile – typically computed from probabilistic genotyping software - can be misleading even if the hypotheses are mutually exclusive and exhaustive. The paper shows that, in contrast to single profile DNA ‘matches’ where the only residual uncertainty is whether a person other than the suspect has the same matching DNA profile, it is possible for all the genotypes of the suspect’s DNA profile to appear at each locus of a DNA mixture, even though none of the contributors has that DNA profile. In the absence of other evidence, the paper shows it is possible to have a very high LR for the hypothesis ‘suspect is included in the mixture’ even though the posterior probability that the suspect is included is very low. Yet, in such cases a forensic expert will generally still report a high LR as ‘strong support for the suspect being a contributor’. 

The problems are especially acute where there are very small amounts of DNA in the mixture (so-called low template DNA abbreviated to  LTDNA). In certain circumstances, the use of the LR may have led lawyers and jurors into grossly overestimating the probative value of a LTDNA mixed profile ‘match’.

Full paper (pdf):

 

Thursday, 17 September 2020

UK: Plotting new Covid cases per 1000 tested

Following on from my analysis of the trend in Covid deaths (and using the same dataset), here is a plot (for each day since 8 April*) of the number of new cases per 1000 people tested. Contrary to what is being shown by the media (see below), this plot is the one that should be used to base decisions on if and when new social distancing or lockdown rules are needed.

When we consider that a higher proportion of new cases now (compared to the first 2 months) are minor or totally asymptomatic**, I find it is rather incredible that the following graph of just new cases (that strongly indicates 'second wave') is the one used by almost all media outlets.  This graph does not take account of the increase in number of people tested. Neither does it take account of the decrease in proportion of deaths per cases, nor the false positive test rate (and so further exagerrates the scale of the problem):


Of course this graph conveniently shows 'strong evidence' of a 'second wave' - something which better fits the narrative of a lot of influential people.


*This is the first date for which there is a record of number of people tested.

**As evidenced by the continued very low death rates shown in my previous piece:


UPDATE 18 Sept. Here are the latest data on hospitalizations and deaths from the Government's own Coronovirus website:



See also



The Cab accident problem: new insights into probabilistic reasoning when there is uncertainty about the witness reliability

 

One of the classic problems used to evaluate how well lay people perform probabilistic updating is the "Blue/Green Cab accident problem" (or equivalently with buses). The problem is usually expressed as follows:

 A cab was involved in a hit-and-run accident at night. Two cab companies, the Green and the Blue, operate in the city. 

90% of the cabs in the city are Green and 10% are Blue. 

A witness identified the cab as Blue. The court tested the reliability of the witness under the circumstances that existed on the night of the accident and concluded that the witness correctly identified each of the two colours 80% of the time and failed 20% of the time. 

What is the probability that the cab involved in the accident was Blue rather than Green? 

A major finding of the original study was that many participants neglected the population base rate data (i.e. that 90% of the cabs are Green) entirely in their final estimate, simply giving the witness’s accuracy (80%) as their answer. 

However, what was not considered in this and similar studies was any uncertainty about the witness reliability (this is called second order uncertainty). If the 80% figure was based on 80 correct answers in 100 then the 80% estimate seems reasonable. But what if it was based only on 5 tests in which the witness was correct in 4? In that case there is much more uncertainty about the "80%" figure; using a Bayesian network model to get the correct 'nomative' solution, it can be shown that while the witness's report does increase the probability of the cab being Blue, it simultaneously decreases our estimate of their future accuracy (because Blue cabs are so uncommon).

A new paper by lead author Stephen Dewitt that addresses how well lay people reason about this second order uncertainty has been published in Frontiers in Psychology. It was based on a study of 131 participants,  who were asked to update their estimates of both the probability the cab involved was Blue, as well as the witness's accuracy, after they claim it was Blue. While some participants responded normatively, most wrongly assumed that one of the probabilities was a certainty; for example, a quarter assumed the cab was Green, and thus the witness was wrong, decreasing their estimate of their accuracy. Half of the participants refused to make any change to the witness reliability estimate. 

Dewitt, S., Fenton, N. E., Liefgreen, A,  & Lagnado, D. A. (2020). "Propensities and second order uncertainty: a modified taxi cab problem". Frontiers in Psychology, https://doi.org/10.3389/fpsyg.2020.503233

Monday, 14 September 2020

COVID19 trend plots: deaths by numbers tested

The UK now enters a new period of 'semi lockdown' with private gatherings restricted to groups of 6 people. This seems to be in response to an 'increase in confirmed cases'; but such an increase is inevitable as more people are being tested (especially young people, many of whom will not suffer any major symptoms), and with the false positive problem. It seems to me that the most relevant COVID19 trend plot should be number of deaths per 100K people tested. I've not seen such a plot  - so I just downloaded the relevant raw data from https://ourworldindata.org/covid-deaths and produced the following plot for the UK from 1 April 2020:

 


As the 'tail' is too small to see, here is the same data, but just from 1 May 2020


For comparison, here is the plot for the USA. Not the same kind of decrease because it is made up of many large states which had increasing infection rates at different times.

 

And another comparison is Israel, which is the first country to implement a second complete shutdown (beginning this week). The issue here is that, because numbers are relatively small and outbreaks occur mainly in very localised orthodox (Jewish and Muslim) communities, bigger variation than UK/USA is inevitable 

 


And here are the three countries together plotted on the same scale. UK now doing much better than USA and Israel (which both are similar now despite being very different in May)

 


 It does seem that the introduction now of the new UK rules is a very strange move..... 


See also https://probabilityandlaw.blogspot.com/search/label/COVID

Sunday, 19 July 2020

A privacy-preserving Bayesian network model for personalised COVID19 risk assessment and contact tracing



Concerns about the practicality and effectiveness of using Contact Tracing Apps (CTA) to reduce the spread of COVID19 have been well documented and, in the UK, led to the abandonment of the NHS CTA shortly after its release in May 2020.

One of the key non-technical obstacles to widespread adoption of CTA has been concerns about privacy. In this new paper our group present a causal probabilistic model (a Bayesian network) that provides the basis for a practical CTA solution that does not compromise privacy. Users of the model can provide as much or little personal information as they wish about relevant risk factors, symptoms, and recent social interactions. The model then provides them feedback about the likelihood of the presence of asymptotic, mild or severe COVID19 (past, present and projected). When the model is embedded in a smartphone app, it can be used to detect new outbreaks in a monitored population and identify outbreak locations as early as possible. For this purpose, the only data needed to be centrally collected is the probability the user has COVID19 and the GPS location.

The paper contains details of how to download and run the model.


And here is a 7-minute video showing the model in action:


Thursday, 16 July 2020

The need for causal models to understand and explain whether statistics provide evidence of racially biased policing


Even before the recent George Floyd case, there has been much debate about the extent to which claims of systemic racism are supported by statistical evidence. Contradictory conclusions have been made about whether unarmed blacks are more likely to be shot by police than unarmed whites using the same data. The problem is that, by relying only on data of ‘police encounters’, there is the possibility that genuine bias can be hidden.

In this short paper we provide a causal Bayesian network model to explain this bias – which is called collider bias or Berkson’s paradox – and show how the different conclusions arise from the same model and data. We also show that causal Bayesian networks provide the ideal formalism for considering alternative hypotheses and explanations of bias.