Skip to content
Insights Usability & UX

How Many Usability Testing Participants Do You Actually Need?

"Five users" is the answer everyone has heard. It is a genuinely useful rule of thumb, and it is also routinely misapplied. Here is what it actually means, when it holds up, and when you need a different number entirely.

Where it comes from

The "five users" rule, and where it comes from

The "five users" figure comes from research by Jakob Nielsen and Tom Landauer in the early 1990s, looking at how many usability problems get found as you add more participants to a study. The pattern they found was a curve of diminishing returns: the first few participants surface most of the major problems, and each additional person finds proportionally less that is new. If you're newer to the method itself, our guide to what usability testing is covers the basics first.

Nielsen's widely cited summary was that testing with five participants in a single round finds around 85% of the usability problems present in an interface, for a single, fairly consistent group of users completing a defined set of tasks. Run a second round of five after making fixes, and you catch most of what is left. That is the basis for the now-standard advice: small rounds, run often, beat one enormous study run once.

Why it works

Why five often works in practice

The logic holds up well in a specific, common situation: you have one main type of user, a focused set of tasks, and you are testing iteratively, fixing issues and testing again as the design evolves. In that context, five participants per round is efficient. You see the same handful of problems crop up repeatedly by participant three or four, diminishing returns set in fast, and your time is better spent fixing what you have found and testing the next iteration than running a sixth or seventh session that mostly confirms what you already know.

This is why five (or a tight 4-5 range) is the basis for our Essentials package: it is the right size for a focused study where the goal is to find and fix the biggest problems in a design, quickly, and move on.

Where it falls short

Where five participants is not enough

You have more than one distinct user group. The "85% with five" figure assumes a reasonably homogeneous group of users with similar goals, contexts, and levels of familiarity with your product. If your product serves clinicians and patients, or new customers and power users, or people using assistive technology and people who are not, those groups can experience completely different problems. Five participants drawn from a mixed pool might miss an issue that affects every single person in one of those groups, simply because nobody from that group happened to be in the sample. The fix is not "more participants from the same pool", it is roughly five per distinct group.

You are testing low-frequency or high-stakes tasks. Some tasks are used by every participant in every session: navigation, search, the core flow. Others are used occasionally, by a minority, but matter enormously when they go wrong, things like cancelling a subscription, requesting a refund, or reporting a safety issue. A small sample can easily miss these entirely, simply because nobody in the room happened to need that task during their session.

You need to make a quantitative claim. "Find the biggest problems" is a different question to "is task completion meaningfully higher with design B than design A?" or "what proportion of users can complete this task unaided?". Comparative or statistical claims need a sample large enough to produce a result that is not just noise, typically 20 or more participants per condition, sometimes considerably more depending on the effect size you are trying to detect. This is the territory of surveys and quantitative research, not a focused qualitative usability round.

You are testing for accessibility and inclusion. People who use assistive technology, screen readers, switch access, voice control, screen magnification, are not a single homogeneous group. Two screen reader users on different software, with different levels of experience, can have very different experiences of the same interface. A handful of sessions with people who use assistive tech is far better than none, but treating it as equivalent to "five users covers it" understates how varied this group's needs and workarounds actually are. We write more about this in our guide to inclusive research recruitment.

How we approach it

How we work out the right number for your study

Rather than starting from a fixed number, we start from the questions you need answered, and work backwards to a sample size and structure that will actually answer them. In practice that comes down to a handful of questions:

How many distinct user groups matter to this study? One group, focused tasks, iterative testing: five is usually right. Two or three meaningfully different groups (by role, by need, by familiarity with the product): we plan for roughly five per group, sometimes run as separate rounds so findings do not get lost in the mix.

Are you trying to find problems, or prove a number? Finding and prioritising usability problems is qualitative work, and a small, well-chosen sample is efficient. Proving that one design measurably outperforms another, or that X% of users can complete a task, is quantitative work, and needs a larger sample run with more rigour around consistency between sessions.

What happens if you miss something? For a marketing site, missing a minor edge case is a minor cost. For a clinical tool, a financial transaction, or a safety-critical workflow, the cost of missing something is much higher, which can justify a larger sample, more diverse recruitment, or sessions specifically targeting edge-case tasks rather than just the common path.

Is this a one-off check, or part of an ongoing process? If you are going to test again after this round, a smaller study now, with a follow-up round after changes, often beats one larger study trying to catch everything in a single pass.

Iteration vs one big study

Two rounds of five usually beats one round of ten

A common instinct, especially when budget approval is hard to get, is to "make it count" by running one large study with as many participants as the budget allows. In most cases this is the wrong instinct. Two rounds of five, with design fixes made in between, will surface more of the problems that actually matter than one round of ten, because the second round is testing a different (better) design, and because seeing the same issue recur after a fix tells you whether your fix actually worked.

The exception is when the design is genuinely close to final and the question has shifted from "what's wrong with this" to "does this work well enough, for enough people, across enough scenarios". At that point, a single larger and more comprehensive round, our Impact package of 8-10 sessions, makes more sense: more segments covered, more edge cases tested, and a report that can support a launch decision rather than the next design iteration.

In practice

Putting it into practice

If you only take one thing from this: do not let "five users" become a reason to skip testing because you "only" have budget for a small study, and do not assume a bigger number automatically means better research. The right sample size depends on how many distinct groups of people use your product, what kind of question you are answering, and what stage your design is at.

When we scope a project, working out the right number of participants, and the right mix of groups, is part of the discovery conversation, not an afterthought. If you are not sure where your study sits on this spectrum, that is a normal place to start from. We will help you work out what "enough" looks like for your specific question, whether that is a focused 4-5 session round, an 8-10 session study across multiple segments, or something else entirely.

Get the free sample size cheat sheet

A one-page reference: how many participants you need by method, how segments multiply the number, and the four questions that settle it for your own study.

We will also add you to our occasional newsletter. Unsubscribe anytime.

We fully recognise the effort your team invested in recruitment, moderation, and analysis, and we genuinely appreciate the quality of the discussions and reporting.

UK Operations Manager Medicsen

Trusted PPI and PPIE delivery partner to the NIHR HealthTech Research Centre in Accelerated Surgical Care.

Reviewing findings on a phone alongside printed research materials

Not sure how many sessions you need?

Tell us about your product, your users, and the decisions you need to make. We will recommend a sample size and structure that fits.

Start a conversation

We usually respond within one working day.

Prefer to talk it through first?

Book a free 20-minute call. No obligation, no sales pitch. We'll tell you honestly whether research is worth it for your decision, and what it would cost.