Skip to content

Real people · Not synthetic data · UK-wide

AI Product Testing: Real Users, Not Synthetic Data

If your product includes a chatbot, AI copilot, recommendation engine, or any other AI powered feature, the people who matter are the real humans who will use it, not a simulated panel or another AI playing the role of your customer. We test AI features with real recruited participants, so you find out how people actually understand, trust, and use your AI before they do it on your live product.

Chatbots & Assistants AI Copilots Recommendation Engines Generative Features Trust & Explainability

Why AI features need real people, not synthetic ones

AI generated personas and simulated users are trained on existing patterns. They can describe what a "typical" user might do, but they cannot show you what happens when a real person with a real task encounters your AI for the first time, gets an answer they do not trust, or asks a question your model was never designed for.

Testing AI with AI also creates a circularity problem. If both the product and the evaluator are generative models, you risk validating your AI against its own blind spots, rather than against the people it is actually built for. We set out the evidence in why AI personas cannot replace real user research.

We bring real recruited participants into contact with your AI powered product and watch what actually happens: where they get value, where they get confused, where they stop trusting the system, and where they need a human to step in.

Real participants
Recruited to match your actual users, not generated personas or AI generated synthetic responses
Trust & explainability
We test whether people understand, trust, and correctly act on what your AI produces
Actionable findings
Prioritised findings with video evidence, covering both UX and AI specific risks

AI features we test

From customer facing chat assistants to AI features quietly working in the background, we design research around how real people experience your AI.

Chatbots & Virtual Assistants

Test how real users phrase questions, react to responses, recover from misunderstandings, and decide whether to trust or escalate to a human.

Conversation design Escalation paths Trust signals

AI Copilots & Writing Tools

Watch how people use AI suggestions inside your product. Do they accept, edit, or ignore them, and does the assistance speed people up or get in the way?

Suggestion UX Editing behaviour Workflow fit

Recommendation Engines

Understand whether recommendations feel relevant, how people react when they are wrong, and whether users notice or care that an algorithm is involved.

Relevance Personalisation User control

Generative Content Features

Test AI generated text, images, or summaries for clarity, accuracy perception, and whether users can tell, or need to know, what was generated.

Content trust Disclosure Quality perception

AI Search & Summarisation

See whether AI generated answers and summaries help people find what they need faster, or introduce new confusion and verification work.

Findability Verification behaviour Source trust

Automated Decision Support

For tools that recommend or make decisions, test whether users understand the basis for a decision, when they would override it, and what evidence they need to feel comfortable.

Explainability Override behaviour Confidence

How AI product testing works

The same rigorous process as our usability testing, with extra attention to how people interact with AI generated output.

1. Discovery call

We discuss what your AI feature does, who uses it, and what "working well" looks like, including any known failure modes or edge cases. Free of charge, usually 30-45 minutes.

2. Research design

We design tasks and scenarios that reflect real use, including situations where the AI might get something wrong, so we can see how people respond.

3. Participant recruitment

We recruit real users matching your audience. For regulated or clinical contexts, this can include patients, carers, or specific professional groups.

4. Testing sessions

Moderated sessions where participants use your AI feature on real or realistic tasks. We probe how they interpret, trust, and act on what the AI produces.

5. Trust & explainability analysis

Beyond standard usability findings, we analyse how participants understood the AI's reasoning, noticed errors, and decided whether to rely on it.

6. Actionable report

Findings covering usability, trust, and AI specific risks, with video evidence and prioritised recommendations your team can act on.

What we look for

Beyond standard usability, AI features raise questions about trust, errors, fairness, and when a human should take over.

Trust calibration

Do people trust the AI the right amount, not too much and not too little?

Best when You are launching a feature where over reliance or under reliance on AI could cause real harm or missed value.

What you get Evidence of how confident users feel, when that confidence is misplaced, and what changes it.

You leave with Recommendations for framing, disclaimers, and confidence indicators that calibrate trust appropriately.

Error recovery

What happens when the AI gets it wrong, and can users tell?

Best when Your AI sometimes produces incorrect, irrelevant, or unhelpful output, as most do.

What you get Evidence of whether users notice errors, how they react, and whether recovery paths work.

You leave with Design recommendations for error states, fallback options, and human handoff points.

Bias & representation

Does the feature work as well for everyone who needs to use it?

Best when Your AI feature serves a diverse audience, including people with different digital confidence, language needs, or accessibility requirements.

What you get Evidence from inclusive recruitment that reflects your real user base, not just confident early adopters.

You leave with Findings highlighting where the experience breaks down for specific groups, and what to fix.

Human handoff

When should the AI step back and let a person take over?

Best when Your product combines AI with human support, whether that is a customer service team, a clinician, or an account manager.

What you get Evidence of where users want or need a human, and whether your current handoff points match that.

You leave with Recommendations for where to add, simplify, or remove escalation points.

Who we work with

From startups building AI native products to established healthcare and regulated organisations adding AI features carefully.

Healthtech & Digital Health SaaS & B2B Platforms NHS & Healthcare Customer Service & Support Tools Fintech & Banking Ecommerce & Retail Education & EdTech Public Sector & Government Digital Pharmaceutical & Life Sciences Startups Building AI Products

AI product testing packages

Transparent pricing, the same as our other usability and UX research services.

Quick Validation

Essentials

From
£6,000

Focused testing of a single AI feature or flow. Right for validating a chatbot, copilot, or recommendation feature before or shortly after launch.

  • 4-5 moderated sessions
  • Screened participant recruitment
  • Trust and explainability analysis
  • Findings presentation with video clips
  • Prioritised recommendations
  • 2-3 week turnaround
Get started

Multi-Phase

Custom Programme

From
£16,000

Ongoing research partnership as your AI features evolve. Test across iterations, combine with PPIE or design validation, or run continuous trust monitoring.

  • Tailored research design
  • Multiple testing rounds across releases
  • Mixed methods available
  • Diverse participant segments
  • Embedded researcher support
  • Flexible delivery timeline
Discuss your needs

All packages include: Participant incentives, session recording and storage, synthesis and analysis covering both usability and AI specific findings, video clip highlights, and clear documentation your team can act on immediately.

Combining services: AI product testing pairs well with design validation for regulated products, or PPIE when AI features are used in healthcare and patient facing contexts.

Customise this estimate for your own study →

Common questions about AI product testing

Why test AI features with real users instead of synthetic personas or AI generated feedback? +
AI generated personas and synthetic user panels can only reflect patterns already present in their training data. They cannot show you how a confused first time user phrases a request to your chatbot, where trust breaks down when an AI gives a wrong answer, or how someone with low digital confidence reacts to a recommendation they do not understand. Real participants surface the unscripted reactions, hesitations, and workarounds that synthetic data cannot produce.
What kinds of AI features can you test? +
We test chatbots and virtual assistants, AI copilots and writing tools, recommendation engines, generative content features, AI powered search and summarisation, and automated decision support tools. Whether the AI sits in the background, such as a recommendation algorithm, or is the main interface, such as a chat assistant, we design research around how real users understand, trust, and act on what the AI produces.
How do you evaluate trust and explainability in AI features? +
We look at whether users understand what the AI is doing, why it produced a particular result, and what to do when it gets something wrong. This includes testing how AI generated content is labelled, whether people notice and act on confidence indicators or disclaimers, how errors are handled, and whether users feel comfortable relying on the output for decisions that matter to them.
How is AI product testing different from standard usability testing? +
The core method, moderated sessions with real recruited users, is the same. What changes is the focus. Alongside standard usability questions like navigation and task completion, we pay close attention to trust calibration, how users interpret AI generated content, whether they notice when the AI is wrong, and how comfortable they are handing over judgement to a system. These behaviours only emerge when real people interact with a real AI feature.
Can you test AI features in healthcare or other regulated contexts? +
Yes. We have extensive experience in NHS, healthtech, and regulated environments, and AI features in these contexts often need particularly careful testing around trust, safety, and clinical or regulatory acceptability. We can combine AI product testing with our PPIE and design validation services to produce evidence suitable for governance and regulatory conversations.
How much does AI product testing cost? +
AI product testing starts from £6,000 for five to six sessions, from £9,500 for eight to ten, and from £16,000 for a multi-phase programme. Recruitment, participant payments, moderation, analysis and a findings report are included. Testing in a regulated or clinical context adds roughly 15 per cent for the extra governance, consent and documentation involved. Build a rough estimate.

We fully recognise the effort your team invested in recruitment, moderation, and analysis, and we genuinely appreciate the quality of the discussions and reporting.

UK Operations Manager Medicsen

Trusted PPI and PPIE delivery partner to the NIHR HealthTech Research Centre in Accelerated Surgical Care.

Reviewing findings on a phone alongside printed research materials

Ready to find out how your AI feature actually performs?

Tell us what you have built and what you are unsure about. We will suggest the right approach, whether that is AI product testing on its own or combined with usability testing, UX research, or PPIE.

Start a conversation

We usually respond within one working day.

Prefer to talk it through first?

Book a free 20-minute call. No obligation, no sales pitch. We'll tell you honestly whether research is worth it for your decision, and what it would cost.