False Needles in Haystacks

👇 Get new sketches each week
Why rare events produce surprising numbers of false positives
If you have a test designed to detect something, a false positive is when your test gives you a positive result when the thing isn’t actually there. Often, we may be looking for something that’s rather rare, say 1 in 1,000 cases. This low base rate can produce a rather unintuitive result: even with what seems like an accurate test, more of your positive results may be false than true.
David Spiegelhalter sums it up in his book The Art of Uncertainty with this analogy:
When you are looking for a needle in a haystack, even if you have good eyesight, most of what looks like needles will be hay.
— David Spiegelhalter
A Worked Example of False Positives
Let me give an illustration. An example he gives in the book concerns facial recognition. Suppose, from pictures in a crowd, a facial recognition system correctly identifies 7 out of 10 people on a watchlist—a 70% detection rate. And this same system has a false-alert rate of just 1 in 1,000, so that when scanning 1,000 people who aren’t on a watchlist, only 1 would come back as a (false) positive.
Now suppose that in a crowd of 10,000 people there are 10 people on the watchlist in the crowd and the facial recognition system scans everyone. Of the 10 people on the watchlist, it will return 7 positive alerts and miss 3 people. Of the remaining 9,990 people in the crowd, at a 1-in-1,000 false-alert rate, the system will trigger about 10 more positive alerts.
So, the system—with a false alert rate of just 1 in 1,000, remember—would return 17 positive alerts, with 7 true positives and about 10 false positives.
For any one individual with a positive alert, there’s therefore a less than 50% chance (7/17=41%) that the person is on the watchlist.
For a system with what seemed a low false-alert rate, it is surprising that a positive alert in this case is still more likely to be false than true.
The difficulty in these cases is that what we are looking for is itself rather rare. So even if a test is good at correctly identifying cases and has a low false-alert rate, when the event happens only very occasionally, false positives can easily overwhelm correct identifications.
For me, the needle-in-a-haystack analogy works very well. We usually imagine the problem is simply finding the needle among all that hay because there’s only one needle and they look rather alike. What I don’t immediately appreciate is that, before finding it, I would find a lot of things that looked like a needle but turned out not to be one.
Other Examples
You may not often find yourself dealing with the results of facial recognition tests on dangerous people, granted. However, this sort of result crops up in other situations. Here are a few:
Healthcare Testing
Suppose a test is developed to screen for a very rare condition. Even if the test is very good at finding the condition and has a low false alert rate, if the condition is sufficiently rare—a very low base rate—we can expect it to falsely identify many people who don’t have the condition.
Accordingly, screening programmes often have a second, more specific diagnostic test after an initial positive result.
Airport Security
Airport security screens enormous numbers of ordinary travellers while genuine threats are, fortunately, extremely rare. As a result, even screening procedures with a very low false-positive rate could flag many regular travellers for every genuine threat it identifies. Perhaps this has happened to you.
Fraud Detection
Electronic payments are really an incredible thing. I worked for some time at a financial institution, and the number of payments and transactions received and sent just by our company was astonishing. Multiply that by the number of financial companies in a city, and I remember thinking that the challenge of regulating and policing fraudulent payments was quite incredible.
If the base rate of fraudulent payments to legitimate payments is very low, then we can expect our fraud detection systems to flag a number of legitimate payments in their search.
Defect Detection
Suppose you run a factory that produces very high-quality components. Only 1 in 10,000 is defective. An automated inspection system that only very occasionally rejects a good component can still be expected to identify many components as defects that may turn out to be fine.
Car Alarms
Suppose a car alarm goes off on your street. What is the likelihood that someone is actually trying to steal a car? Only once have I actually seen someone breaking in (in Rome) for all the car alarms I’ve ever heard. So I pretty much assume an alarm is a false positive.
A car alarm is looking for a rare event, yet it spends thousands of hours monitoring just in case. It’s possible no one will try to break into your car in its lifetime. Even an alarm with a low false-alert rate could therefore sound more often because of minor disturbances than because someone is actually trying to break into the car.
When there's a lot of hay and very few needles, expect false positives. How rare is the thing you're looking for?
--
I had the pleasure of speaking with Sir David on a very entertaining and informative episode of the Sketchplanations podcast .
Related Ideas to False Needles in Haystacks
Also see:

