Early Signals vs. Large-Scale Experiments

Early Signals vs. Large-Scale Experiments

Chapter Ten: Testing Assumptions, Not Ideas

Inevitably, someone on your team is going to raise a concern with making decisions based on small numbers. How can we have confidence in the data if we talk to only five customers? You might be tempted to test with larger pools of people to help get buy-in. But this strategy comes at a cost—it takes more time. We don’t want to invest the time, energy, and effort into an experiment if we don’t even have an early signal that we are on the right track.

Rather than starting with a large-scale experiment (e.g., surveying hundreds of customers, launching a production-quality A/B test, worrying about representative samples), we want to start small. You’ll be pleasantly surprised by how much you can learn from getting feedback from a handful of customers.

Imagine we test the assumption, “Our subscribers want to watch sports,” as described above. We show participants a mockup of our “home screen,” and we ask them what they’d like to watch. Four out of ten choose a sporting event, soaring past the threshold we set for our success criteria. What will we do next?

It depends. We’ve made this assumption more known. However, we can’t conclude it’s true. It still carries risk. The question becomes, “How much risk?” If we have another assumption on our assumption map that is now riskier, we want to switch gears and test that assumption. But if this assumption continues to be our riskiest assumption, and it carries more risk than our organization can stomach, then we need to continue to test it. We need to start defining the next-level experiment that will allow us to collect more data.

Perhaps, as a next step, we decide to add a section to our real “home screen” promoting an upcoming sporting event. When users select it, it informs them that we are considering adding sports to our lineup, and we ask for feedback by way of a thumbs-up or a thumbs-down. We also give them the option to submit comments. We think we can get this experiment live with a week of development work. We decide to collect data from 500 participants (which we think we can do in 3 days), and we set our success criteria to at least 100 of the 500 participants giving us a thumbs-up. This is a classic smoke-screen test.

So why didn’t we start here? Our first test was designed to be completed in a day or two. This test will take up to two weeks—maybe longer, if we need to get permission from stakeholders and/or need to wait for an upcoming release cycle. We don’t want to invest this time, energy, and effort until we’ve received an early signal that we are on the right track.

With assumption testing, most of our learning comes from failed tests. That’s when we learn that something we thought was true might not be. Small tests give us a chance to fail sooner. Failing faster is what allows us to quickly move on to the next assumption, idea, or opportunity. Karl Popper, a renowned 20th-century philosopher of science, in the opening quote argues, “Good tests kill flawed theories,” preventing us from investing where there is little reward, and “we remain alive to guess again,” giving us another chance to get it right.

As we test assumptions, we want to start small and iterate our way to bigger, more reliable, more sound tests, only after each previous round provides an indicator that continuing to invest is worth our effort. We stop testing when we’ve removed enough risk and/or the effort to run the next test is so great that it makes more sense to simply build the idea.