Asking about what happened, not what would
Everything you have written so far is a hypothesis, and it is a hypothesis assembled by people who already know the answer. This is where it meets somebody who does not.
Three questions destroy customer research, and all three feel natural.
"Would you buy this?" Hypothetical, and the direction of the error is known. In experiments comparing stated to real valuations, people overstate, reliably. The size varies enormously between studies and between goods, from near zero to a factor of several, which is exactly why you cannot correct for it. Purchase intention does predict purchase to a degree: across 40 studies covering more than 65,000 consumers and 200 products, the mean correlation between stated intention and actual purchase was 0.49, with a range from −0.13 to 0.99. That correlation is respectable by social science standards and unusable as a forecast, and it is weakest for new products, which is the case people use it for.
"What features would you like?" It asks a buyer to do your job, and they will oblige. The answer is a list of things that sound reasonable, generated in seconds, with no cost attached to any of them.
"What do you think of this advert?" It converts a buyer into a critic. The register shifts, and criticism is what you get, calibrated to the person's idea of what a marketer wants to hear.
What replaces all three is reconstruction of a purchase that already happened.
From there, the sequence follows the buyer rather than your list. The questions that do work are all about events.
- When did you first notice this was a problem? What happened that day?
- What did you try first?
- Who else had to be involved?
- What made you start looking for something else?
- What did you type, or who did you ask?
- What nearly stopped you?
- What did you decide against, and why?
- How long between first looking and deciding?
Ask about a specific past event, not a general policy. "What do you usually do" produces a self-description. "What did you do in March" produces a fact, with a date and a sequence and the parts the person had forgotten they had told you.
Two things about how you ask that are measured well enough to be worth changing your behaviour over.
Wording moves aggregate answers by amounts that would embarrass most researchers. The classic demonstration asked logically equivalent questions and found 54 per cent said yes to forbidding something while 25 per cent said yes to allowing it, a 21 point gap; a meta-analysis of subsequent forbid and allow experiments put the mean asymmetry at 14 percentage points. That literature is about attitude questions on social topics rather than purchases, so take the mechanism rather than the number: the level of an answer is a property of the question as much as of the person. This is also the defence of the practice, and it holds: comparisons between people asked the same question are far more robust than the absolute level of any single answer.
Leading questions do something worse than skew the answer. In the founding demonstration, participants who were asked how fast cars were going when they "smashed" into each other estimated 40.8 miles per hour against 34.0 for "hit", in a 45-person experiment, and a week later 32 per cent of the "smashed" group, in a second experiment of 150 people with 50 per condition, reported seeing broken glass in a film that contained none, against 12 per cent of controls. That is students and a film clip, not a customer interview, and it is the mechanism behind why a discovery interview that supplies the buyer's motives will hear them repeated back a fortnight later as memory.
One practical note on mode. A meta-analysis of 61 studies found that people distort answers slightly less to a screen than to a human interviewer, at d = −0.19. That is a small effect and it is a reason to use a form for the awkward question about price, not a reason to stop interviewing. Interviews get you the sequence, which no form does.
Three interviews is the minimum this course accepts and it is a floor, not a target. Three will not give you a distribution. Three will reliably give you one thing you did not know, which is usually a trigger you had never considered or a hesitation nobody in your organisation had heard of, and that is the return on the week.
A buyer tells you they would definitely pay more for a version with priority support. What is the correct weight to put on this?
You are interviewing a customer who bought six months ago. Which question is most likely to produce something you did not already know?
Your three interviews produce three different triggers. What does this tell you?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.