Improving public understanding of probability and risk with special emphasis on its application to the law. Why Bayes theorem and Bayesian networks are needed
I was recently watching a re-run of the classic 1978 Michael Cimino film “The Deer Hunter”. It contains one of the most iconic scenes in cinema history involving a ‘game’ of Russian roulette forcibly played by two American soldiers held captive in Vietnam. Although I have seen the film several times, this scene never seems to lose its impact. As I am currently teaching a new course on Risk Assessment and Decision Making, it also occurred to me that the scene provides a rich source of examples to illustrate core concepts of probability and risk including: probability and odds, basic probability axioms, conditional probability, risk and utility, absolute versus relative risk, event trees, and Bayesian networks.
So, I have written a short paper which hopefully has something of value both for people with no background in probability/statistics and also people who do, but want to find out more:
Researchers at UCL and Birkbeck have published an important study on the benefits of using a Bayesian Network (BN) tool to solve the kinds of complex problems that intelligence analysts are confronted with.
Example of the type of problem considered. Participants had to answer questions such as which group was most likely responsible for the attack based on various details about multiple informant sources and their accuracy
The study provides strong empirical evidence that if you provide basic training to use the BARD tool for constructing BNs then this
improves the ability of individuals to solve complex probabilistic
reasoning problems, compared to a control group receiving only generic
training in probabilistic reasoning.
The full details of the paper (which includes a link to all of the problems and data) are:
Cruz, N., Desai, S. C., Dewitt, S., Hahn, U., Lagnado, D., Liefgreen, A., Phillips, K., Pilditch, T., and Tešić, M. (2020). "Widening Access to Bayesian Problem Solving". Frontiers in Psychology, 11, 660. https://doi.org/10.3389/fpsyg.2020.00660
*I declare an interest here: the BARD tool was developed from the AgenaRisk API.
** Again I declare an interest: I was involved with some of the training
Suppose we wanted to estimate how many car owners there are in the UK and how many of those own a Ford Fiesta, but we only have data on those people who visited Ford car showrooms in the last year. If 10% of the showroom visitors owned a Fiesta, then, because of the bias in the sample, this would certainly overestimate the proportion of Ford Fiesta owners in the country.
Estimating death rates for people with COVID-19 is currently undertaken largely along the same lines. In the UK, for example, almost all testing of COVID-19 is performed on people already hospitalised with COVID-19 symptoms. At the time of writing, there are 29,474 confirmed COVID-19 cases (analogous to car owners visiting a showroom) of whom 2,352 have died (Ford Fiesta owners who visited a showroom). But it misses out all the people with mild or no symptoms.
Read more:
COVID-19 tests: how they work and what's in development
Concluding that the death rate from COVID-19 is on average 8% (2,352 out of 29,474) ignores the many people with COVID-19 who are not hospitalised and have not died (analogous to car owners who did not visit a Ford showroom and who do not own a Ford Fiesta). It is therefore equivalent to making the mistake of concluding that 10% of all car owners own a Fiesta.
There are many prominent examples of this sort of conclusion. The Oxford COVID-19 Evidence Service have undertaken a thorough statistical analysis. They acknowledge potential selection bias, and add confidence intervals showing how big the error may be for the (potentially highly misleading) proportion of deaths among confirmed COVID-19 patients.
They note various factors that can result in wide national differences – for example the UK’s 8% (mean) “death rate” is very high compared to Germany’s 0.74%. These factors include different demographics, for example the number of elderly in a population, as well as how deaths are reported. For example, in some countries everybody who dies after having been diagnosed with COVID-19 is recorded as a COVID-19 death, even if the disease was not the actual cause, while other people may die from the virus without actually having been diagnosed with COVID-19.
However, the models fail to incorporate explicit causal explanations in their modelling that might enable us to make more meaningful inferences from the available data, including data on virus testing.
What a causal model would look like.Author provided
We have developed an initial prototype “causal model” whose structure is shown in the figure above. The links between the named variables in a model like this show how they are dependent on each other. These links, along with other unknown variables, are captured as probabilities. As data are entered for specific, known variables, all of the unknown variable probabilities are updated using a method called Bayesian inference. The model shows that the COVID-19 death rate is as much a function of sampling methods, testing and reporting, as it is determined by the underlying rate of infection in a vulnerable population.
Therefore, different countries may appear to have different death rates, but only because they have applied different sampling and reporting policies. It is not necessarily because they are managing the virus any better or that the virus has infected fewer or more people.
With a causal model that explains the process by which the data is generated, we can better account for these differences between countries. We can also more accurately learn the underlying true population infection and death rates from the observed data. Such a model could be extended to include demographic factors, as well as social distancing and other prevention policies. We have developed such models for many similar problems and are currently gathering data required for populating the kind of model that we outline in the above figure.
Random testing
In the absence of community-wide testing, only random testing applied throughout the population will enable us to learn about the number of people with COVID-19 who are asymptomatic or have already recovered. Only when we know how many people don’t show symptoms, will we know the underlying infection and death rate. It will also enable us to learn about the accuracy of the tests (false positive and false negative rates).
Random testing therefore remains the most effective strategy to avoid selection bias and reduce the distortions in reported statistics. Ideally, this should be combined with a causal model.
Random testing would be ideal.SamaraHeisz5/Shutterstock
Currently it seems there are no state-wide protocols in place in any country for randomised community testing of citizens for COVID-19. Spain did attempt it. But that involved purchasing large volumes of rapid COVID-19 tests, and they soon discovered that some Chinese-sourced tests had poor validity and reliability delivering only 30% accuracy – resulting in high numbers of false positives.
Read more:
COVID-19 tests: how they work and what's in development
Countries like Norway have proposed introducing such tests, but there is uncertainty around how to legislatively compel citizens to test – and what might constitute an appropriate randomisation protocol. In Iceland, they have voluntary sampling which has covered 3% of the population, but this isn’t random. Some countries with large scale testing, like South Korea, might get closer to being random.
The reason it is so hard to achieve random testing is that you have to account for several practical and psychological factors. How does one collect samples randomly? Gathering samples from volunteers may not be sufficient as it does not prevent self-selection bias.
During the H1N1 influenza pandemic of 2009–2010, there was a lot of anxiety about the disease that created “mass psychogenic illness”. This is when hypersensitivity to particular symptoms leads to healthy people self-diagnosing as having a virus – meaning they would be highly incentivised to get tested. This could, in part, further contribute to false positive rates if the sensitivity and specificity of the tests are not fully understood.
While self-selection bias is not going to be eliminated, it could be reduced by running field tests. This could involve asking the public to volunteer samples in locations where, even in a lockdown state, they might be expected to attend and also from those in self-imposed isolation or quarantine.
In any event, it is important to note that when statistics are communicated at press conferences or in the media, it is very important that their limitations are explained and any relevance to the individual or population are properly delineated. It is this which we contend is lacking in the current crisis.
We know some strawberry flavoured sweets are contaminated. But, if wrapper colour is not a reliable indicator of the flavour of sweet it contains, what do we learn about the proportions of strawberry and contaminated sweets if we only test sweets with red wrappers?
The current COVID-19 strategic testing strategies - implemented to inform policy making - focus primarily on people already hospitalized with significant symptoms or on people most at risk. This seems to make sense for short-term medical reasons, but such testing is highly biased with sub-optimal consequences. Without understanding the causal explanations for the resulting data from such testing we end up with highly misleading conclusions about infection and death rates. Starting with an analogy of testing sweets for contamination, this short paper illuminates the need for random testing combined with causal models:
I've written about Simpson's paradox before. Given its importance in highlighting the need for causal explanations of observed data, I've been using it to motivate students on my new module on risk assessment and decision analysis for data science. I have put together a couple of videos with examples to explain it graphically (see below). I wanted to base one of the videos on the example of 'exercise v cholesterol' presented in the excellent "Book of Why" by Pearl and Mackenzie:
But it turns out there is a problem with this example. It assumes that in the 'data' in the real world (the left hand figure), older people are the ones who do most exercise. This is clearly not the case. At first I thought this was due to a simple ‘typo’ in that they labelled the age groups the wrong way round (i.e. the 10 – 20 – 30 – 40 – 50 age groups should be reversed). But if you reverse them you hit a different error – this time it would show that older people have lower cholesterol than young people, which is again clearly wrong.
So whichever way you spin this, the example simply does not make sense in the ‘real world’ because it does not make sense for the chosen attributes.
However, the example can be 'fixed' by considering instead 'exercise v junk food consumption' because - in the real world - it is the case that older people not only exercise less than younger people but they also eat less junk food. (**22 March 2020 UPDATE)
I have prepared a (6-minute) video using this example:
And here is another video (5-minutes) explaining a more common example of Simpson's paradox:
**22 March 2020 update: It seems that in some age categories there might be a problem also with my assumption. My colleague Marko Tesic points out:
I’m just wondering about the relationship between exercise and junk food intake within each age group. I’m not sure the association is negative for each age group (although I do think that when considering the whole population this association is positive). People who exercise often eat quite a lot: the more people are active the more fuel their body needs to recover, in particular if people want to gain muscle weight (which is often the case with young people).
Now, it’s not unlikely that a bunch of the food that people who excise eat is actually junk food. The attached paper suggests exactly that. Namely, they find that many people, in particular young people, indulge in junk food after exercise. So I think that the association between exercise and junk food intake may not be negative within each age group. Rather, it’s perhaps positive for teenagers and young adults, close to no association for mature adults and negative for pensioners. This is still interesting as it’d be showing that the general population association does not hold in all age groups and that there’s a partial (rather than complete) reversal in the association.
Simone Dohle, Brian Wansink, and Lorena Zehnder (2014). Exercise and Food Compensation: Exploring
Diet-related Beliefs and Behaviors of Regular Exercisers.Journal of Physical Activity and Health doi:http://dx.doi.org/10.1123/jpah.2013-0383
Yesterday the PM announced the importance of not cancelling major sporting events; and the Premier League announced there would certainly be no cancellations of this weekend's matches. But anybody with any football knowledge knew that - whatever the Government's risk mitigation plans were for Coronovirus - they would become irrelevant if a single high profile Premiership player or manager became infected.
As soon as it was confirmed that Arsenal manager Mikel Arteta had the virus last night, it was inevitable that a total shutdown of all professional football would start and that is precisely what has happened. Again, anybody with any UK football knowledge knows that this is also a game-changer as far as the whole UK economy is concerned. Millions who were previously unmoved to make any changes will voluntarily go into lock-down after panic buying (see the above immediate response).
Note that nothing much changed when it was announced that the Government's own Health Minister got the virus a few days ago. But one key football person getting it was the single trigger for mass change.
I can only assume that the Government risk experts/advisors did not include a single person with football knowledge.....
p.s. from a purely selfish perspective, as a Spurs fan, I am delighted that - by the time the Premiership resumes - we might have some of these players fit again..