- Research splits two ways: generative vs evaluative, and qualitative vs quantitative.
- Sample-size "rules" are mostly myths; recruiting is the real bottleneck and analysis is the work.
- Even without a dedicated researcher, a senior PM owns the research discipline.
The phrase "we did some user research" is doing a lot of work in product teams, and not in a good way. In the same breath it covers a properly recruited diary study with twelve participants, a Google Form sent to the email list, a single Zoom call with one customer the sales lead happened to like, and the cross-functional workshop where the team voted on coloured Post-it notes about what users probably want. Those are not the same thing. Treating them as the same thing is how product organisations end up with shelves of decks labelled "Insights" and a roadmap that ignores them.
Research is a discipline. It has methods. The methods have failure modes. Picking the wrong method for the question wastes calendar weeks and produces evidence that looks convincing but isn't. Picking the right one, then executing it badly (bad recruiting, leading questions, no analysis) produces the same outcome with extra steps. The senior product person's job is to know which method answers which question, what the method costs to run properly, and what it will actually tell you when it lands.
Generative versus evaluative.
The first cut is the most important and the one most teams skip. Generative research opens up the problem space. You don't yet know what to build, or whether anything needs building. You're looking for unmet needs, undocumented behaviours, jobs the user is hiring some other product to do. The output is a problem framing, not a feature list. Evaluative research closes a question. You have a thing (a prototype, a flow, a live feature) and you want to know whether it works. The output is a verdict and a fix list.
Most product teams over-index on evaluative and under-invest in generative. Usability tests are visible. The team watches them, the screenshots make it into the deck, so they get funded. Generative interviews are quieter, slower, and produce conclusions that take weeks to land. So generative gets squeezed, the team ships features that nobody asked for, and the post-mortem reports "good execution, weak product-market fit". The phrase is a tell. The product was answering a question nobody asked.
The 2×2 that actually matters.
Seven methods cover the field for most product teams. 1-on-1 interviews are the workhorse of generative qualitative: sixty to ninety minutes, semi-structured, recorded and transcribed. Diary studies capture behaviour over time, where memory is unreliable (anything to do with workflow, routines, or emotion). Contextual inquiry means watching users in their environment. It's invaluable when the user's world matters, useless when the artefact is a flat web tool with no context to observe. Surveys scale qualitative hypotheses into quantitative shape: never the first method to reach for, often the right one to validate a generative finding before betting on it.
Usability testing is the headline evaluative method: five to eight users, scripted tasks, observation, think-aloud. A/B tests are the gold standard for causal measurement of changes to a live product, covered in detail in N°35. Analytics straddles both columns, used generatively to surface opportunity (where do users drop off, where do they stall) and evaluatively to measure the impact of a change.
The sample-size lies.
The most-cited number in user research is Jakob Nielsen's "five users find eighty-five per cent of usability problems". It's true for a single round of usability testing on a single artefact within a single user persona. It is not true for anything else. It is not true for generative interviews (where you're trying to surface a range of unmet needs, not converge on one). It is not true for surveys (where your sample needs to support whatever subgroup analysis you'll do). It is not true for A/B tests (where statistical power dictates the floor). And it is not true for usability testing if you have multiple personas — five users per persona, not five users total.
This single misapplication of one rule of thumb has done more damage to research practice than any other. Senior practitioners hear "five is enough" and treat it as the universal answer. Two months later the team has shipped a feature based on five interviews with one persona out of three, and is surprised when retention in the other two cohorts drops.
Recruiting is the bottleneck.
Recruiting is the unglamorous engine of a research practice. You cannot run weekly research if you don't have a recruiting pipeline. Most product teams discover this the first time they try. By week three the sales team has stopped forwarding candidates, the customer success team's patience has run out, and the team is sending desperate Slack messages asking who-knows-someone-who.
Three solutions. Build a panel — a maintained list of customers and prospects who've opted in to be contacted for research, with attributes you can segment on. Pay a recruiting platform: User Interviews, Respondent.io, dscout. Useful for B2C and for B2B when you need participants outside your customer base. Standardise the incentive: a published rate per session, paid promptly, no negotiation. (Sixty to a hundred pounds for sixty minutes is the going rate for most B2C in the UK; double that for specialist B2B.)
The "we did a focus group" anti-pattern deserves a paragraph of its own. Focus groups are the worst common method in product research. The dominant participant shapes the consensus; the quiet participant says nothing; the moderator's nodding signals which answers earn approval. Nielsen Norman Group has written about this for twenty years and the field still won't listen. If you find yourself running a focus group because it's faster than six 1-on-1s, run the six 1-on-1s. The data is dramatically better and the cost is comparable.
Analysis is the work.
Most teams stop after the interviews are done. They have a folder of transcripts and a slide that says "Key themes". The slide was written from memory, after a single pass through the notes, by the person who ran the interviews — who is, by definition, biased toward the things they noticed in the moment. That isn't analysis. That's recall dressed up as insight.
Proper analysis means coding transcripts (tagging quotes with themes), then doing affinity mapping (clustering the codes into emergent groupings), then refining the groupings against the transcripts to make sure they hold up. It's tedious. It takes roughly the same time as the interviews themselves did. Tools like Dovetail, Reduct and Marvin have made it faster than it used to be, but they haven't eliminated the discipline. Skipping analysis is how you get a research debrief that confirms whatever the lead PM already believed.
ResearchOps as a function.
At enough scale, research stops being something a PM does on the side and becomes a function in its own right. ResearchOps is the supporting infrastructure — the panel, the recruiting workflow, the incentive payment system, the repository of past studies, the templates, the consent forms, the calendar logistics. It is to user research what DevOps is to engineering: not the work itself, but the system that lets the work happen weekly without grinding to a halt.
If your organisation has more than three product teams and you don't have at least one ResearchOps lead, every PM is running their own recruiting and templating from scratch, which means most of them aren't doing it at all. Centralising the ops makes individual research cheaper, not more expensive. That's the case to take to the CFO.
The senior PM's responsibility, even without a researcher.
Plenty of product teams don't have a dedicated researcher. That doesn't excuse the senior product person from the discipline. It means the discipline is theirs. The senior PM picks the method, runs the sessions properly, codes the transcripts, runs the affinity map, presents the findings. The output may be less polished than a researcher's would be, but the rigour must not be. Outsourcing rigour to "we didn't have a researcher" is the most common failure mode in mid-market product organisations, and it's the one that compounds. Every quarter of skipped research locks in the wrong assumptions for the quarter after.
How we use this at Product Pieces.
When a partner brings us a product whose roadmap is stuck on "we don't know what to build", the diagnostic usually finds the gap is research. Not "they don't have a researcher" (they often do) but that the practice has collapsed into evaluative-only, or that recruiting has stalled, or that analysis is being skipped. The Discovery Set Piece rebuilds the practice in four to six weeks: pipeline, methods, templates, a backlog of generative studies, and a handover plan to keep it running without us.
Next week's issue picks up the thread. Once you have a research practice, how do you make it weekly rather than quarterly? That's Teresa Torres's question, and the answer is the Opportunity Solution Tree.