Understanding False Positives and False Negatives

Understanding False Positives and False Negatives

Chapter Ten: Testing Assumptions, Not Ideas

Now, this method isn’t flawless. When working with small numbers, we will encounter false positives and false negatives. Let’s explore the impact of these errors on our work.

In your first round of experimenting, it is possible that you’ll select 10 participants who all hate sports. We can mitigate the risk of this by choosing a variety of folks. In other words, we don’t want to choose 10 participants from Honolulu, Hawaii (where no major sports teams reside), and expect to get reliable results. Instead, we want to select for variation in geographic location, demographics, TV-watching behavior, etc., as best we can. However, even if we select for variation, it is still possible that none of our participants like sports when our larger population does. That’s because we aren’t doing the work to select a representative sample, and we aren’t testing with large-enough numbers.

When this happens, when our experiment fails, even though our larger population exhibits the behavior that we want to see, we call this a “false negative.” Our test is providing data that indicates our assumption is faulty when it may not be.

But what’s the cost of this false negative? In this particular example, where our assumption is testing our target opportunity “Our subscribers want to watch sports,” we might consider abandoning the opportunity. However, we aren’t likely to make this decision based on one failed test. Instead, if we are running tests across our set of ideas, we will have additional data points to help us evaluate the target opportunity.

For example, if we test assumptions across three different ideas, all exploring if our subscribers are interested in sports, and all of them fail, then the chance that all of them are false negatives goes down. More likely, we’ll get conflicting results. We’ll see one assumption fail and another one pass. We’ll need to dig in to learn why. In the worst-case scenario, one of our results will be a false negative, and we’ll have to run additional experiments to evaluate our assumption. However, if our tests are small, this costs us only a day or two. This isn’t a very costly false negative.

Most of our assumptions, however, aren’t testing the opportunity. They are testing some aspect of a particular solution. When these assumptions fail, we typically design around them. We evolve our ideas so that they no longer depend on the faulty assumption. For example, if we are testing the assumption “Our subscribers know where to find sports on our platform,” and it turns out to be problematic, we can always redesign the interface to make sports easier to find. If our failure was a false negative, it’s possible we might redesign our interface when we don’t need to. But if further testing shows that our iteration works, the cost of this false negative is only the time it took to do the redesign. Again, this false negative isn’t that costly, as long as we keep our iterations and our future testing small.

Additionally, when we run fast iterations, we are in a better position to make decisions using multiple data points from several tests rather than make decisions based on a single data point. We can test if our subscribers want to watch sports through a number of testing methods. We can ask them about their past viewing behavior and see if they have watched sports in the past. We can show them a mockup and ask them what they would like to watch right now. We can simulate the moment before the big game and see if they choose our service. Instead of throwing out an assumption based on one data point, we can draw conclusions from the set of assumption tests. Researchers call this triangulation.52 It’s using a mix of research methods to better understand the assumption we are testing.

Finally, even in the worst-case scenario, when we do decide to abandon an idea or an opportunity and it turns out it was based on a false negative, it’s still okay. There are hundreds, if not thousands, of ideas that could address our target opportunity or opportunities that could drive our desired outcome. When we throw one away needlessly, it’s not that costly, as long as we find an idea that does work or an opportunity that does have an impact. Remember, there isn’t one right idea or one right opportunity. We can afford false negatives because ideas and opportunities are abundant.

Now let’s turn to false positives. A false positive is when our test gives us data suggesting that our assumption is true, when it isn’t. This sounds far riskier than a false negative, but, in practice, it’s not. Suppose we run our small test, and we learn that everyone wants to watch sports, so we call our test a success, and we move forward. Remember, we aren’t making a go/no-go decision based on one assumption test. We are either moving on to test another assumption related to the same idea, or we are running a bigger, more reliable test on the same assumption. If our idea really is faulty, odds are that our next round of assumption testing will catch it. False positives usually get surfaced in successive rounds of testing. The cost of a false positive in a small test is usually the time and effort required to run the next-bigger test. That’s not trivial, but we still avoid the far-bigger cost of building the wrong solutions.

I want to be clear: There is a cost to false negatives and false positives. And we should be aware that these costs exist. But the cost is not so great that we should be starting with large-scale, quantitative experiments every time. If we did that, we would never ship any software. Our tests would simply take too long. The vast majority of the time, you will learn plenty from your small-scale tests.