Simulate an Experience, Evaluate Behavior

Simulate an Experience, Evaluate Behavior

Chapter Ten: Testing Assumptions, Not Ideas

With assumption testing, our goal is to collect data that will help us move the assumption from the right to the left on our assumption map (see Chapter 9)—we are starting with an assumption that has weak supporting evidence, and our goal is to collect more evidence. Just like with interviewing (see Chapter 5), to collect reliable data, we want to focus on collecting data about what people actually do in a particular context, not just what they think or say they do in general.

A strong assumption test simulates an experience, giving your participant the opportunity to behave either in accordance with your assumption or not. This behavior is what allows us to evaluate our assumption.

To construct a good assumption test, you’ll want to think carefully about the right moment to simulate. You don’t want to simulate any more than you need to. This is what allows you to iterate quickly through several assumption tests.

Let’s return to our streaming-entertainment example, where we are trying to address the target opportunity “I want to watch sports” and we’ve brainstormed a set of solutions—adding local channels, licensing events directly from sports leagues, and bundling our service with a sports provider.

If we are testing the assumption, “Our subscribers want to watch sports,” we might simulate the moment when someone is browsing their streaming options, trying to decide what to watch. We could do this by mocking up the “home screen” that users see when they turn on the streaming-entertainment service. We could present them with several options, including a handful of sporting events, popular TV shows, and recent movie releases. We could then ask them, “What would you prefer to watch right now?”

If we are trying to test the assumption “Our subscriber wants to watch sports on our platform,” we might simulate the moment in time when the big game is about to start. Since our platform doesn’t currently offer sports, we can’t simply ask them which service they prefer. Instead, we might present them with three subscription services (including ours), tell them that the game is available on all three services, and ask them to choose a service to stream the game.

Now, neither of these simulations is perfect. What I say I want to watch when talking to a stranger might differ from what I want to watch when I’m at home by myself. Or I might favor one subscription on one day and another subscription on another day. That’s okay. Perfect simulations are hard to come by. Instead, we’ll account for these shortcomings when we decide how to evaluate the results of our simulation.

Notice how all three ideas depend on the assumption “Our subscriber wants to watch sports.” This is an assumption that is core to the target opportunity, so if this assumption is false, we can abandon our set of ideas. However, only the first and second ideas depend on the assumption “Our subscribers want to watch sports on our platform.” It’s common for ideas to share assumptions. It’s one of the reasons why assumption testing is faster than idea testing. Assumption tests don’t merely give us a go/no-go decision for an individual idea; they help us evaluate sets of ideas. We’ll also see later how shared assumptions will help us generate even better ideas after a round or two of testing.

Once we’ve identified the moment in time that we want to simulate, we now need to define how we’ll evaluate the behavior we observe. Our goal when considering our simulation evaluation is to define what success looks like. In other words, if our assumption is true, what would we expect the participant to do?

For example, when observing people selecting what to watch, we might evaluate how many people choose to watch a sporting event vs. how many people don’t. If our assumption is true—that our subscribers do want to watch sports—we would expect at least some of them to choose a sporting event in our simulation.

When simulating the moment before the big game starts, we might want to evaluate how many people think of our platform as the place to watch the game vs. another platform. If our assumption is true—that our subscribers did want to watch sports on our platform—then we would expect some of them to choose our service over the competitors’.

Now, in both simulations, we say “some people” should exhibit the behavior we expect. The problem with “some” is that your product manager might define “some” as 5 out of 10, and your designer might define “some” as 20 out of 100.

If you run your simulation with 10 people, and 6 people choose sports, your product manager is going to think your assumption is now more known, and your designer is still going to be skeptical. The challenge with this scenario is that, as a team, you didn’t learn anything new from this assumption test because you disagree on what the results mean.

To avoid this situation, we want to get specific with our evaluation criteria. Instead of saying, “Some people choose sports,” we want to say, “At least 3 out of 10 people choose sports.” We want to define both how many people we’ll test with and how many people need to exhibit the behavior that we expect to see.

By defining these criteria upfront, you are doing two things. First, you are aligning as a team around what success looks like so that you all know how to interpret the results. This will help to ensure that your assumption tests are actionable. And second, you are helping to guard against confirmation bias. Remember, confirmation bias makes us more likely to see the evidence that our idea will succeed than the evidence that it might not succeed. If we don’t define our success criteria upfront, when we try to interpret the results, our brains will actively look for evidence that supports the assumption, and we’ll likely miss the evidence that could refute it. To avoid this, we want to define what success looks like upfront (before we see the results).

So how do we choose the numbers? This is a subjective decision. Your goal is to find the right balance between speed of testing and what aligns your team around an actionable outcome. You want to test your assumption with as few people as possible (as it will be faster) but with the number of people that still gives your team the information they need to act on the data. Now remember, you aren’t trying to prove that this assumption is true. The burden of truth is too much. You are simply trying to reduce risk. Keep your assumption map in mind. Your goal is to move the assumption from right to left. How many people would convince you this assumption is more known? That’s the negotiation you are having as a team.

If your simulation is less than optimal, as we saw with the above examples, you’ll need to modify your numbers to accommodate for these shortcomings. If someone raises the concern that some sports fans might want to watch sports on our platform, but in the moment we ask them, they might be more likely to choose a comedy, then you might lower your threshold for success to account for that. If someone else is worried that choosing from three subscriptions biases the results in favor of your subscription (because the reality is people have, on average, five subscription services), then you could either decide to change your mockup to show five services or raise your threshold for your success criteria.

The key outcome with this exercise is to agree as a team on the smallest assumption test you can design that still gets you results that the team will feel comfortable acting on.