Instrument Your Evaluation Criteria

Instrument Your Evaluation Criteria

Chapter Eleven: Measuring Impact

Start by instrumenting what you need to collect to evaluate your assumption tests. As you build your live prototypes54, consider what you need to measure to support your evaluation criteria. Don’t worry about measuring too much beyond that. For example, in the story that opens this chapter, we had several assumptions we needed to test:

  • Students will start more searches if we ask them easier questions.
  • Students will view jobs that we recommend.
  • Students will apply to jobs that we recommend.

We defined evaluation criteria for each assumption:

  • 250 out of 500 visitors will start their search using our new interface. (Remember, we were seeing only 180 out of 500, or 36%, start their search on our old interface. We wanted to see a big jump in search starts to warrant such a different interface.)
  • At least 63 of our 500 students will view at least one job. (Our current interface was performing at 81 out of 500. We set our initial criteria lower, because we knew our canned searches weren’t perfect, and we were confident we could improve them over time.)
  • At least 7 of our 500 students would apply for a job. (Our current interface was performing at 12 out of 500. Again, we set our initial criteria lower because we knew our results would get better over time.)

With this evaluation criteria in mind, here’s what we measured:

  • # of people who visited the search start page
  • # of people who started a search
  • # of people who viewed at least one job
  • # of people who applied for at least one job

Notice how we are counting the number of people who took a specific action and not counting the number of actions. This is an important distinction to pay attention to when instrumenting your product. Sometimes you’ll want to count people. Other times you’ll want to count actions. A good way to suss this out is to ask, “If one person did many actions, does that create as much value as many people doing one action?” If you need many people to take action, you’ll want to count people. If it doesn’t matter how many people take action, you’ll want to count actions.

In this case, our assumptions were more about the perception of our new interface. We were concerned that students might not trust our recommendations. So, we wanted to measure how many people engaged with our job listings. We wanted to make sure that the new interface was working for more people than the old interface.

However, when we started to measure the relevance of our saved searches, we started to count actions. We wanted to know how many jobs people found to be compelling. This wasn’t a straightforward metric. If someone views 25 jobs, it might be because they are finding 25 jobs that interest them. Or it might be because it took 25 tries before a job interested them. For relevance, we took two measurements. We measured the position of a job view in the search result (e.g., a student viewed the first vs. third job in the search results). We also measured the ratio of job views to job applications (e.g., the number of jobs someone had to view before they applied for a job).

Counting people helped us understand how many of our students were having success on our platform. Counting actions helped us understand how hard each student had to work to find success.

Notice, however, that we did not start by measuring everything. We didn’t track every click on every page. We started with our assumptions, and we measured exactly what we needed to test our assumptions.