Skip to content

Explainer · Methods · How it works

What Is Usability Testing?

Usability testing is a research method where real people attempt realistic tasks with a product while a researcher watches, to find where they get stuck, confused or give up before launch. This guide covers how it works, the main types and methods, when to use it, and how it differs from UX research, A/B testing and a UX audit.

The definition

Usability testing is a research method in which real people attempt to complete tasks on a website, app, or digital product while a researcher observes and records what happens. The goal is to identify where users get confused, stuck, or give up: the specific problems that analytics cannot explain and that internal teams, too close to the product, typically cannot see.

It answers a deceptively simple question: can people actually use this thing? Not whether it is technically functional, not whether it looks good in a prototype review, but whether real users with real goals can navigate it, understand it, and get what they came for.

Usability testing is distinct from user research that explores what people need. It evaluates whether a product meets those needs in practice. It is typically conducted with a prototype or live product, though even paper sketches and wireframes can be tested with real users early in the design process.

Real users
The only reliable way to find usability problems is to watch real people use the product, not ask colleagues to review it
5 users
The research-backed minimum to surface the majority of critical usability issues in a single user group
Specific and actionable
Usability testing produces findings your design and development team can act on immediately, not vague recommendations

Types of usability testing

The main distinction is between moderated and unmoderated testing. From there, sessions can be remote or in-person, and the format adapts to what you are testing.

Moderated usability testing

A researcher runs each session live, guiding participants through tasks and asking follow-up questions in real time. When a user does something unexpected, the researcher can explore why. This produces richer, more nuanced insight than automated testing because the researcher can probe what is happening beneath the surface, not just record that it happened. Moderated testing is the standard for significant design decisions, complex products, and any situation where understanding the reasoning behind user behaviour matters.

Unmoderated usability testing

Sessions run automatically through a software platform without a live researcher. Participants complete tasks and think aloud, and recordings are reviewed afterwards. Faster and cheaper per session than moderated testing, and useful for collecting quick signal on simple, well-defined questions. The limitation is that unmoderated testing cannot explore unexpected behaviour in real time: it records what happened, not why. Best used for validation tasks with clear success criteria, not for understanding complex or nuanced problems.

Remote usability testing

Sessions conducted via video call (Zoom, Teams, or a specialist platform), with participants anywhere in the UK. Remote testing gives fast access to a wider participant pool, lower cost per session, and is well-suited to most digital products. Participants use the product in their own environment, which can reveal context that a lab setting would mask. Remote sessions can be moderated or unmoderated.

In-person usability testing

Researcher and participant are in the same physical space: at your office, a research facility, or the participant's own environment. In-person testing picks up body language, physical context, and environmental factors that remote sessions cannot capture. It is better suited to hardware and medical device testing, complex B2B tools, participants who are not comfortable with video calls, and situations where stakeholder observation is important for buy-in.

Formative usability testing

Testing conducted during the design process, typically on prototypes or early versions of a product, with the aim of identifying problems while they are still easy and cheap to fix. Formative testing is diagnostic: the goal is to find issues and understand them well enough to solve them. It does not need to be statistically rigorous; the focus is on uncovering problems, not measuring performance.

Summative usability testing

Testing conducted at the end of a design phase or on a finished product to measure usability against defined criteria: task completion rates, error rates, time on task, satisfaction scores. Summative testing is evaluative rather than diagnostic. It is often used for regulatory evidence (medical devices, regulated software), benchmarking against a previous version, or comparing two competing design directions with measurable outcomes.

How usability testing works

The structure of a usability testing project is consistent regardless of what you are testing or whether sessions are remote or in-person.

1. Define the research questions

What do you need to find out?

Usability testing works best when the questions are specific. Not "is this usable?" but: can users find the pricing page? Do users understand what the onboarding flow is asking them to do? Where do people drop off in the checkout? Sharper questions produce more useful findings.

2. Design the protocol

Tasks, scenarios, and participant criteria.

The session guide sets out the tasks participants will attempt, written as realistic scenarios rather than instructions. Recruiting criteria define who the participants need to be: the characteristics, behaviours, and demographics that define your actual user group.

3. Recruit participants

Real users who match your audience.

Participants are screened against your criteria, incentivised, and scheduled. The quality of the recruiting determines the quality of the findings: participants who do not resemble your actual users will not surface the problems your actual users experience.

4. Run the sessions

Moderated, recorded, observed.

Each session (typically 45-60 minutes) involves one participant attempting the tasks while the researcher observes, prompts, and probes. Sessions are recorded. Stakeholders can watch live or review recordings. The researcher asks participants to think aloud as they work through tasks.

5. Analyse and synthesise

Patterns across sessions.

The researcher reviews recordings and notes across all sessions, identifying patterns: issues that appeared repeatedly, points of confusion multiple participants shared, tasks that consistently took too long or produced errors. Findings are categorised by severity.

6. Report and recommend

Prioritised, actionable, evidenced.

The findings report presents issues prioritised by severity, supported by video evidence, with clear recommendations. Critical problems first, then medium friction, then quick wins. Written for both stakeholders who need the headline and designers who need the detail.

How usability testing compares to other research methods

Usability testing sits within a wider toolkit of user research methods. Knowing what it does and does not do well helps you choose the right method for the question.

vs analytics

Analytics tells you what happened: where users dropped off, which pages they visited, which buttons they clicked. Usability testing tells you why: what confused them, what they misunderstood, what they were trying to do when they left. Analytics surfaces the symptom; usability testing finds the cause. The two are complementary: analytics identifies where to look, usability testing explains what is happening there.

vs A/B testing

A/B testing measures which of two versions performs better against a metric (conversions, sign-ups, time on page). It tells you that B is better than A, but not why, and not what would make C even better. Usability testing produces insight that can generate and inform hypotheses for A/B testing. The methods work well together: usability testing to identify and understand problems, A/B testing to measure the impact of fixes at scale.

vs user interviews

User interviews explore attitudes, needs, motivations, and mental models: they surface what people think, feel, and want. Usability testing observes behaviour: what people actually do when they try to use a product, which is often very different from what they say they would do. Interviews are best for discovery and understanding; usability testing is best for evaluation. Both are valuable; neither replaces the other.

vs surveys

Surveys measure self-reported attitudes and satisfaction at scale. They are good for tracking sentiment across a broad user base, or for measuring satisfaction after a product change. They cannot reveal usability problems, because users often cannot identify or articulate what made something hard to use. Surveys measure opinions; usability testing observes behaviour.

vs heuristic evaluation

Heuristic evaluation (a UX audit) involves an expert reviewing a product against established usability principles. It is faster and cheaper than user testing and can identify a range of potential issues without recruiting participants. The limitation is that expert reviews reflect what an expert thinks users will find difficult, not what real users actually find difficult. Heuristic evaluation is useful for quick triage; usability testing provides empirical evidence from the actual user population.

vs focus groups

Focus groups bring multiple people together to discuss a topic, generating attitudes and reactions through group conversation. They are not suitable for evaluating usability: social dynamics in a group setting distort individual behaviour, and discussing a product is not the same as trying to use it. For evaluating digital products, usability testing with individual participants produces more reliable and useful evidence than group discussions.

When to use usability testing

Usability testing is relevant at almost every stage of a digital product's life. The question is what you are testing and what you need to find out.

Early in design

Test paper sketches, wireframes, or low-fidelity prototypes. The goal at this stage is to find out whether the core concept and structure make sense to users before you invest in detailed design. Problems found here are cheap to fix.

Pre-launch

Test a high-fidelity prototype or near-final build before you ship. This is where the majority of usability testing happens: finding problems before real users encounter them in production, when the cost of fixing them is still low and the reputational cost of poor usability has not yet landed.

Post-launch optimisation

Test a live product where analytics have flagged a drop-off point or conversion problem you cannot explain. Usability testing turns the analytics question ("why are people leaving here?") into a specific, actionable answer.

After a major redesign

Redesigns introduce new patterns and flows that existing users have to relearn. Testing with both new and existing users after a redesign surfaces the specific things that have become harder and validates that the new design actually works better.

For regulatory or compliance evidence

Medical devices, regulated software, and healthcare digital services often require documented evidence of summative usability testing. Structured task-based testing with defined pass/fail criteria produces the human factors evidence needed for regulatory submissions.

When launch has stalled

If conversion rates are lower than expected, customer support tickets are high, or activation rates are not improving, usability testing provides a faster path to a diagnosis than iterating on guesses. A few sessions with real users typically reveals what is actually happening.

Common questions about usability testing

What is the difference between moderated and unmoderated usability testing?

Moderated testing involves a live researcher who can probe unexpected behaviour and explore the reasons behind what users do. Unmoderated testing runs automatically through software and records what happens without a researcher present. Moderated testing produces richer insight; unmoderated testing is faster and cheaper for straightforward validation tasks.

How many participants do you need for usability testing?

Five participants is the research-backed minimum for finding the majority of critical usability issues in a single user group. For studies covering multiple distinct user segments, device types, or accessibility needs, 8-10 or more sessions are appropriate. More is not always better: adding participants without adding diversity rarely adds new findings. See our usability testing service page for how we advise on participant numbers.

Is usability testing the same as user acceptance testing (UAT)?

No. User acceptance testing (UAT) checks whether software meets its requirements: does this feature work as specified? Usability testing checks whether real users can actually use the product effectively. UAT is a functional QA process; usability testing is a research process. A product can pass UAT and still be very difficult for users to navigate.

Can you test a product that is not built yet?

Yes. Figma prototypes, InVision designs, clickable wireframes, and even paper sketches can all be tested with real users. Early-stage testing often produces the most valuable findings, because identifying problems before development begins is far cheaper than fixing them after build.

What makes a usability test reliable?

Well-written task scenarios that reflect what real users are actually trying to do, rather than leading participants towards the expected path. Participants who genuinely represent the target user group. A moderator who observes without directing, probing what they see rather than confirming what they expected. And enough sessions to identify patterns rather than treating individual behaviour as representative.

How does usability testing work for healthcare or NHS products?

The principles are the same, but the participant recruitment is more specialised (patients, carers, clinicians), ethical approvals may apply, and the products often need to work for users with lower digital literacy or access needs. We work alongside PPIE programmes and understand NHS digital service standards. Regulated medical devices require summative testing against defined human factors criteria; see our design validation service for that context.

What is the difference between remote and in-person usability testing?

Remote usability testing uses video call software to run sessions with participants wherever they are. It gives faster access to a wider participant pool, lower cost per session, and is suitable for most digital products. In-person testing brings the researcher and participant together physically, at your office, a research facility, or the participant's own environment. In-person is better for hardware and medical device testing, digitally excluded users, complex B2B tools, and situations where physical context or direct observation by stakeholders matters.

What is the difference between usability testing and UX research?

Usability testing evaluates an existing product or prototype: it measures how well users can complete tasks and identifies where they struggle. UX research is a broader category that includes discovery methods such as user interviews, diary studies, and contextual inquiry, which explore user needs, motivations, and behaviours before or alongside product development. Usability testing asks 'can people use this?'; UX research asks 'what do people need and why?' The two methods are complementary: discovery research to understand the problem, usability testing to evaluate the solution.

How much does usability testing cost in the UK?

A focused study with 4-5 remote participants, including participant incentives, recruitment, moderation, analysis and a findings report, starts from £5,500. A comprehensive study with 8-10 participants runs from £9,500. In-person sessions add venue and travel costs. Multi-phase research programmes start from £16,000.

How long does usability testing take?

A focused study with 4-5 remote participants can be delivered in 2-3 weeks from brief to findings report. A comprehensive study with 8-10 participants typically takes 3-4 weeks. The main variables are participant recruitment time, prototype readiness, and session scheduling. Quick validation studies can be turned around in under two weeks when needed.

What is the difference between usability testing and A/B testing?

A/B testing measures which of two versions performs better with live traffic: it tells you that version B converts better, but not why. Usability testing shows you exactly where users get confused, what they misunderstand, and why they make the choices they do. The two methods work well in combination: usability testing to identify and understand problems, A/B testing to measure the impact of fixes at scale once you know what to change.

We fully recognise the effort your team invested in recruitment, moderation, and analysis, and we genuinely appreciate the quality of the discussions and reporting.

UK Operations Manager Medicsen

Trusted PPI and PPIE delivery partner to the NIHR HealthTech Research Centre in Accelerated Surgical Care.

Reviewing findings on a phone alongside printed research materials

Ready to run usability testing on your product?

Tell us what you are building and what questions you need answered. We will suggest the right approach, scope, and turnaround for your project.

See our usability testing service

Or get in touch directly to discuss your project. We usually respond within one working day.

Prefer to talk it through first?

Book a free 20-minute call. No obligation, no sales pitch. We'll tell you honestly whether research is worth it for your decision, and what it would cost.