JDKRUEGER&COAcademySign inDE
Module 1 of 6 · UX & User Research for E-Commerce

Usability Testing for Shops

⏱ 30 min · After this module, you'll be able to choose the right usability test type for your shop (moderated, unmoderated, mobile), write goal-oriented tasks that reveal real behavior instead of opinions, explain the rule of thumb of around five testers, and prioritize findings by frequency, severity, and proximity to conversion.
← UX & User Research for E-Commerce Usability Testing for Shops 1 / 27
continuing in 5
Start

Usability Testing for Shops

Why real users tell you more about your shop than any analytics chart, and how to use that systematically without a big budget or an external agency.

Assumption → Observation → Finding → Action
Transcript of this slide

Welcome to the first module of the UX Research track. Over the past few years, I've watched shops pour tens of thousands of euros into traffic and then stumble over small surface-level issues. Usability testing is the fastest tool for finding exactly those spots. You don't need a big budget or an external team. You need the right method, the right tasks, and a clear evaluation framework. That's exactly what you'll learn here.

Learning objective

What You'll Learn in This Module

You'll distinguish between moderated, unmoderated, and mobile usability tests by use case.

  • You'll write tasks that generate real behavior instead of opinions.
  • You'll scale tester numbers sensibly and justify that decision internally.
  • You'll prioritize findings by frequency, severity, and business impact.
1
Choose a test type
2
Define tasks
3
Scale testers
4
Prioritize findings
Transcript of this slide

After this module, you'll have four concrete skills. First, you'll choose the right test type: moderated for depth, unmoderated for scale, mobile for the most underestimated lever. Second, you'll write tasks that reveal where users actually get stuck. Third, you'll justify your tester count in a way your team or agency can follow. And fourth, you'll evaluate results not by who's loudest, but by revenue impact. That's the difference between an academic usability report and a conversion decision.

Self-check

Quick Self-Check

When did you last watch a real customer move through your shop, not your team, not your agency?

  • Which three spots in your shop would you test immediately if you knew the results would be actionable?
Key Points

Quick Self-Check

  • 1 When did you last watch a real customer move through your shop, not your team, not your agency?
  • 2 Which three spots in your shop would you test immediately if you knew the results would be actionable?
Transcript of this slide

Before we get into the method, two quick honest questions. Most shop decision-makers have never watched a real user navigate their shop. They look at analytics, heatmaps, conversion rates. But the actual behavior stays invisible. If you answer these two questions for yourself, you already know where your biggest usability levers are.

What this means for you

What Does This Mean for Your Shop?

A single usability test with five users can surface barriers you'd never spot in a thousand rows of analytics data.

  • Most shops don't have usability problems. They have an observation problem.
  • Regular testing helps you avoid costly wrong turns and speeds up your learning cycle.
Assumptions vs. Observed Friction
Transcript of this slide

The business case is straightforward. I've seen shops spend months debating a redesign, then run a single usability test with five users and discover the problem was somewhere else entirely. Most shops don't have usability problems in the classic sense. They have an observation problem. Looking regularly saves you expensive detours and gets you to faster decisions.

Concept

What Usability Testing Actually Measures

Usability measures whether users reach their goals, not whether they like the design.

  • Task Success Rate: the share of testers who complete a task successfully.
  • Time-on-Task, error rate, and perceived difficulty round out the picture.
  • The System Usability Scale gives you a quantitative benchmark score between zero and one hundred.
24 47 71 94 85 Find product 45 Cart 62 Checkout
Example Task Success Rate by Step
Transcript of this slide

The key difference from pure design feedback is that usability testing measures behavior. The central metric is the Task Success Rate. Secondary metrics include Time-on-Task, error rate, and perceived difficulty. If you want, you can add the System Usability Scale. But be careful: numbers alone don't tell you why something goes wrong. That's why observation always needs to go alongside the quantitative data.

Concept

Moderated vs. Unmoderated Tests

Moderated tests are conducted live: you can ask follow-up questions, read body language, and probe spontaneously.

  • Unmoderated tests run independently, scale faster, and cost less.
  • For shops with low traffic or complex decision-making processes, moderated testing delivers more depth per tester.
Moderated: depth + context. Unmoderated: volume + speed.
Transcript of this slide

The first decision you face is moderated or unmoderated. Moderated means live facilitation with follow-up questions and nonverbal signals. It's ideal when the problem area is still unclear. Unmoderated runs independently and scales faster. For shops with low traffic, moderated is often the better choice because each tester delivers more insight. For shops with clear hypotheses and high traffic, unmoderated is more efficient.

Concept

Remote, mobile, or on-site?

Remote tests save travel time and allow for real devices and everyday situations.

  • Mobile tests should always be run on real smartphones, not on a desktop.
  • On-site tests are the right choice for complex products or when you want to observe physical reactions.
18 36 54 72 65 Remote 25 Mobile Remote 10 On-site
Typical distribution of usability tests in e-commerce
Transcript of this slide

The second decision is about location. Remote tests are the standard today: they save time, allow real devices, and show the user in their natural environment. This is especially important for mobile shops, because a smartphone sitting on a desk behaves differently than one used on the go. On-site tests make sense when you sell physical products or want to observe very emotional reactions. For purely digital shops, remote is sufficient and more efficient in most cases.

Concept

The Magic of the Number Five

Jakob Nielsen and Tom Landauer showed that five testers find about 85% of critical usability problems.

  • Every additional tester delivers significantly diminishing returns per euro and hour.
  • It's better to run three iterative rounds of five testers than a single session with fifteen testers and no follow-up.
24 49 73 97 31 1 55 2 70 3 80 4 85 5 88 6
Cumulative problem discovery by number of testers
Transcript of this slide

One key insight from usability research: five testers find about 85% of critical problems. Every additional tester contributes less that's new. The better approach is to observe five testers, fix the most serious issues, then test with five testers again. That kind of iteration delivers more value than a single large test with fifteen or twenty people.

Example

Case Study: The Furniture Shop and the Search

A furniture shop observes five testers searching for a dining table for six people, budget under €800.

  • Three out of five users search for terms like "table six people" or "dining table 180 cm" instead of using the shop's categories.
  • The internal categorization by style doesn't match customers' mental models.
  • Result: Task success rate in search rises from 0.5 to 0.8 after the adjustment.
Before: 50% success. After: 80% success.
Transcript of this slide

A furniture shop tests its product search with five users. Three don't search by style; they search for 'table for six people.' The navigation was organized by style. After the adjustment, the task success rate rose from 50% to 80%. Five users were enough to identify an expensive structural problem. That shows how quickly a focused test can create more value than weeks of discussion.

Concept

Mobile Usability Testing: the Biggest Lever

60 to 70% of e-commerce traffic comes from mobile devices.

  • The mobile conversion rate is typically 30 to 50% lower than the desktop rate.
  • A mobile usability test usually uncovers bigger barriers than a desktop test.
1 3 4 5 3.8 Desktop Conversion (%)DesktopConversion (… 2.1 Mobile Conversion (%)MobileConversion (…
Typical conversion gap: desktop vs. mobile
Transcript of this slide

If you can only run one test, make it mobile. 60 to 70% of your visitors come via smartphone, and the mobile conversion rate in many shops is well below the desktop rate. That's usually not a traffic problem; it's a surface problem: buttons that are too small, nested menus, long forms, or unclear shipping costs. A mobile usability test shows you within an hour exactly where things break down.

Concept

RITE: Test, Fix, Repeat

RITE stands for Rapid Iterative Testing and Evaluation: test today, fix tomorrow, test again.

  • Instead of one large final report, you get a fast learning cycle with direct action steps.
  • This method works especially well for critical paths like search, cart, and checkout.
1
Test with 5 users
2
Fix the most obvious problems
3
Test again with 5 users
4
Prioritize remaining problems
Transcript of this slide

RITE is a pragmatic framework that fits e-commerce perfectly. You test with five users, fix the most obvious problems right away, then test again. That's far more effective than a months-long usability project ending in a thick final report. For critical paths like search, cart, and checkout, this short learning cycle really pays off. It stops you from shipping changes based on assumptions that nobody actually needed.

Scenario

Scenario: Checkout drops off especially hard on mobile

Analytics shows 70% drop-off in the checkout, particularly on mobile devices.

  • A mobile usability test with five users reveals that the dropdown for the delivery address overlaps the payment icon.
  • Three out of five users accidentally tap 'Back' instead of 'Next.'
  • Removing the overlap reduces mobile checkout drop-off by 8 percentage points.
100 mobile checkout visits → 70 abandonments → after fix 62 abandonments
Transcript of this slide

Analytics flags high mobile checkout abandonment, but gives no clue why. A test with five mobile users reveals a dropdown overlapping the payment icon. Users accidentally tap Back. After the fix, abandonment drops by eight percentage points. For a shop with five thousand mobile checkout visits per month and an average order value of €80, that adds up to €40,000 in additional revenue per year. For a single small fix.

Concept

Writing good tasks

Good tasks describe a goal, not the steps to get there: 'Find a dining table for six people under €800.'

  • Bad tasks ask for opinions: 'What do you think of our design?'
  • Every task needs a clear success criterion you can observe.
Goal description vs. step-by-step instructions
Transcript of this slide

The quality of a usability test is determined by its tasks. Good tasks describe a goal and leave the path open. Bad tasks ask for opinions. Every good task has a clear success criterion: product found, added to cart, or checkout completed. When you give tasks like these to your agency or team, you get results you can actually use.

Concept

Prioritizing findings by business impact

Not every problem you find carries the same weight. Evaluate by frequency, severity, and proximity to conversion.

  • A problem in checkout almost always takes higher priority than a problem on the homepage.
  • A frequency-times-severity matrix shows which findings need to be addressed right away.
3 5 8 10 9 Checkout:high 7 Cart: high 4 Search:medium 2 Homepage: low
Example prioritization by business impact
Transcript of this slide

Usability tests quickly produce a lot of findings. Rate each one by frequency, severity, and proximity to conversion. A checkout problem almost always takes priority over a homepage problem. A frequency-times-severity matrix shows immediately which findings should be tackled first. That way you avoid your team spending time on minor cosmetic issues while checkout is on fire.

Example

Before and after: a checkout fix in numbers

Starting point: five thousand mobile checkout visits per month, 70% abandonment, one hundred fifty purchases.

  • Usability test reveals: users can't find the coupon button and suspect hidden costs.
  • After the change: checkout abandonment drops to 65%, purchases rise to one hundred seventy-five.
  • At an average order value of €80, that's €22,400 in additional revenue per year.
4620 9240 13860 18480 14400 Before(annual) 16800 After(annual)
Annual revenue gain from one checkout fix
Transcript of this slide

A shop with five thousand mobile checkout visits per month has 70% abandonment. A test shows that users can't find the coupon button and suspect hidden costs. After the change, abandonment drops to 65%, which means twenty-five additional purchases per month. At an average order value of €80, that's €22,400 in additional revenue per year. That's the economic value of a single well-prioritized usability finding.

Concept

When do you need more than five testers?

When you have multiple distinct user groups, you need three to five testers per group.

  • For quantitative metrics like conversion rate or task success rate, you need twenty to thirty testers per variant.
  • For pure problem discovery, five testers are enough, as long as the group is homogeneous.
8 17 25 33 5 Qualitativeproblems 15 Three usergroups 30 Quantitativemetrics
Recommended number of testers by goal
Transcript of this slide

Five testers is a rule of thumb, not a law. If you have multiple distinct user groups, say B2B buyers and end consumers, you need at least three to five testers per group. If you want to measure quantitative metrics like a task success rate with statistical reliability, you need twenty to thirty testers per variant. But for pure problem discovery, which is what most shops actually need, five testers per homogeneous group is enough.

Common misconception

Common mistakes in usability testing

Mistake one: 'We'll test once the shop is finished.' Tests are most valuable in the early stages.

  • Mistake two: 'We ask users what they want.' Observe behavior, not wishes.
  • Mistake three: 'Five testers aren't enough.' For qualitative problems, five to eight testers are usually sufficient.
  • Mistake four: 'We only test desktop.' Mobile checkout is often the bigger opportunity.
Assumed truths vs. methodological reality
Transcript of this slide

Four common mistakes. First: tests get delayed until everything is done. Second: teams ask for wishes instead of observing behavior. Third: people assume they need dozens of testers. Fourth: only desktop gets tested. All of this prevents real problems from being caught early. Avoid these mistakes and you're already ahead of most competitors.

Exercise

Your quick exercise: the first test task

Pick a core task in your shop, such as the product finder, the shopping cart, or the mobile checkout.

  • Write a task description that defines a real goal without prescribing how to reach it.
  • Note three observation criteria: task success, time taken, and visible friction.
1
Choose a task
2
Define a goal
3
Set criteria
Transcript of this slide

Do this now. Choose a task that's central to your shop. Frame it as a goal, not a set of directions. Then define three observation criteria: Does the user complete the task? How long does it take? Where do they show friction, hesitation, or frustration? This short exercise is the raw material for your first professional usability test.

Concept

Usability Testing in the Research Mix

Usability tests reveal behavior, surveys explain motivation, A/B tests confirm the solution.

  • No method replaces another. Only in combination do they make decisions truly reliable.
  • Use tests to generate hypotheses, not to prove statistical significance.
1
Test: Behavior
2
Survey: Why
3
A/B testing: Solution
Transcript of this slide

Usability testing is one part of a broader research mix. It shows you the behavior. Surveys explain the motivation behind it. And A/B tests confirm whether a measure actually works. No method replaces another. If you only run tests, you miss the why. If you only run surveys, you miss what people actually do. Combine both, and you make significantly better decisions.

Interim check

Quick Check

Usability testing measures behavior, not opinion.

  • Five to eight testers will surface most of the serious problems.
  • Mobile tests and checkout tests have the highest business impact.
  • Findings must be prioritized by frequency, severity, and proximity to revenue.
1
Measure behavior
2
Five testers
3
Test on mobile
4
Prioritize
Transcript of this slide

A quick check before we wrap up. Usability testing measures behavior, not opinion. Five to eight testers will surface most of the serious problems. Mobile tests and checkout tests typically have the highest business impact. And findings need to be prioritized, otherwise you end up fixing cosmetic details while your checkout is on fire.

Summary

Summary

Usability testing is the fastest way to find real friction points in your shop.

  • Choose moderated testing for depth, unmoderated for scale, and mobile for the biggest impact.
  • Good tasks describe goals, not steps, and deliver observable success criteria.
  • Prioritize findings by business impact, not by who speaks loudest.
1
Choose a test type
2
Define tasks
3
Five to eight testers
4
Prioritize and act
Transcript of this slide

The core in four sentences. Usability testing is the fastest tool for finding real friction points. Choose the test type that fits your goal. Write tasks that describe objectives. And evaluate the results based on what helps revenue the most. That's the difference between a nice report and a conversion decision. In the next module, we'll look at how to filter the relevant moments out of thousands of session recordings.

Quiz

Quiz

Test your knowledge.

A shop observes five mobile users during checkout. Three out of five accidentally tap 'Back' because a dropdown overlaps the payment icon. What is the best way for the shop to act on this finding?

Which statement about the task success rate is correct?

You sell high-priced B2B products and don't yet know exactly where new buyers are getting stuck in your shop. Which test format is the better choice to start with?

According to Nielsen and Landauer, approximately what share of serious usability problems do five testers find?

One finding points to a problem on the homepage. Another points to a problem in the mobile checkout. Both occur equally often. How do you prioritize?

Exercise

Exercise

Apply what you have learned right away.

  • 1
    Your personal usability test brief
    mini-audit · approx. 25 min
    Pick a key path in your shop (for example, product search, cart, or mobile checkout). Write three goal-oriented tasks that generate real behavior without pointing users toward the answer. For each task, define a success criterion, an observation criterion, and an estimated business impact. Decide which task should be your top priority and explain your reasoning in terms of frequency, severity, and how close it sits to the conversion.
  • 2
    Evaluate your tester count and test type
    benchmark · approx. 15 min
    Take stock of your current situation: which user groups shop with you, and how different is their buying behavior? Choose between moderated and unmoderated, mobile and desktop, and explain your reasoning. How many testers per group would be right for your next goal?
Reflection

Reflection

A quick look back before you continue.

  • Which core task in your shop (e.g., search, product selection, checkout) would you test first with five users - and why that one specifically?
  • How would you explain to a stakeholder why five mobile testers uncover most critical issues, rather than waiting for a large, expensive test panel?
  • What barrier is currently preventing usability tests from happening regularly in your company (say, following the RITE principle), and how would you remove it?
  • Which decision in your shop could be improved most quickly with a usability test involving five users?
Feedback

Feedback

Was this module helpful for your shop?

Sources

Sources & further reading

Here you will find links and materials to explore the topic in more depth. Take your time.

Overview & learning objective

This module is aimed at shop owners.

After this module, you'll be able to choose the right usability test type for your shop (moderated, unmoderated, mobile), write goal-oriented tasks that reveal real behavior instead of opinions, explain the rule of thumb of around five testers, and prioritize findings by frequency, severity, and proximity to conversion.

Usability Testing for Shops