Intended path observing wrong turn FINDING 5/6 users navigated to Settings before finding Reports — label ambiguity is the root cause usability-testing · assumed path vs actual path · observation reveals what review cannot

Usability Testing

Watch real users attempt real tasks. Then fix what breaks. Usability testing is observing what happens when an unprepared person encounters your design for the first time — not asking what they think, but watching what they do.

User ResearchDesign ValidationNavigation TestingOnboarding EvaluationForm TestingCheckout Optimisation

Two sentences.

Usability testing is the practice of observing real users attempting to complete real tasks with your product — not asking them what they like, but watching what happens when an unprepared person encounters your design and tries to use it. The output is not a satisfaction score; it is a set of observed task failures, hesitations, and navigation errors that reveal exactly where the design is breaking down.

The method was formalised by Jakob Nielsen and Rolf Molich in their 1990 research and has been a cornerstone of UX practice since. No amount of internal critique or expert walkthrough can replicate the moment when someone who was not in the room when a design was made encounters it for the first time. That moment is the only reliable test of whether the design works.

Apply this when…

A new flow has been designed but never been observed in use by someone outside the team
A conversion metric has been declining and the team needs to diagnose the cause
A redesign needs validation that it does not break current task completion behaviour
Onboarding is underperforming and you need to observe where new users lose the thread
Two competing directions exist and user observation can resolve the choice with evidence

When NOT to apply it

Skip it when the question is "which version performs better at scale" (use A/B testing), when you need statistical significance (5-8 participants is directional, not representative), when the work is incremental visual refinement, or when observation would change behaviour (testing a meditation app while being watched).

The mechanism

Usability testing creates a controlled observation condition: a real user, a real task, and no help from the team. Every moment a user hesitates, backtracks, misreads a label, or takes the wrong path is data. The test measures the design's clarity, not the user's ability.

01
Five users reveal 85% of usability problems
Nielsen's 1993 analysis established that five participants identify approximately 85% of usability problems. This means usability testing does not require large samples, and running one round with five users and iterating is more valuable than one large study with twenty. Each user after the fifth adds progressively less new information.
02
Task scenarios, not feature demonstrations
A task scenario must describe a goal without using interface labels. "Find the export function" tells users where to look. "You need to send this report to your manager" describes a goal and lets the interface succeed or fail. If any word in the scenario appears as a label in the interface, rewrite it.
03
What users say and what users do are different data
Behavioural data — where the user clicked, where they paused, what path they took — is more reliable than self-reported data. Users rationalise after the fact and give socially acceptable answers. Think-aloud protocols capture reasoning, but observed paths should be weighted more heavily when the two conflict.
04
Task completion, time on task, and error count
The three primary metrics: did the user complete the task, how long did it take, and how many wrong paths were taken. For diagnostics, error location is most actionable — it identifies exactly which element causes failures. Time on task reveals efficiency problems that completion rate alone misses.

Not the same as user acceptance testing

UAT checks whether the product works as specified — whether engineering matches design intent. Usability testing checks whether the design intent works for users. A product can pass UAT perfectly and fail usability testing comprehensively. Both are necessary and neither substitutes for the other.

GOV.UK's continuous testing and the plain language revolution

When GOV.UK was redesigned starting in 2011, usability testing was not a phase — it was continuous. Every two weeks, the team tested with five to eight members of the public attempting government transactions. Findings fed directly into the next sprint.

The team discovered that jargon — "statutory," "liable," "obligations" — caused consistent failure not because users did not understand the concepts, but because the words triggered uncertainty that broke confidence and stalled navigation. The plain language guidelines that emerged were not written from principles — they were extracted from hundreds of hours of watching real people get stuck on real words.

GOV.UK · Government Digital Service · 2011
Continuous testing produces systematic plain language standards
Sprint cycleBuildTestIterate5 users · 2 hours · every sprintLabel comparisonStatutory obligationWhat you must do by lawPlain language extracted from observed task failureGOV.UK · continuous usability testing · jargon failure → plain language standards
Continuous testing → systematic language standards

Test yourself & see real examples

No examples yet — be the first.

Spotted a flow that clearly has never been tested by a real user — or one so considered it must have been tested extensively? Submit what you observed.

✓ Reviewed before publishing✓ Your name on every example you submit✓ Violation or fix — both welcome

Seen Usability Testing skipped in a real product? Help grow the evidence base.

Where teams go wrong

Asking users what they would do instead of watching what they do. "What do you think of this navigation?" is not a usability task. "Find where you would change your notification settings" is. Self-reported preferences are unreliable; observed behaviour is not.
Writing task scenarios that contain interface labels. "Click the Reports tab" tells the user where to navigate and removes the test. Task scenarios must use goal language. If any word in the scenario appears as a label in the interface, rewrite it.
Testing too late to act on the findings. Post-build testing produces findings that require rework. Prototype-stage testing produces findings that require iteration. The highest-value testing happens before engineering begins, when changing the design costs a day rather than a sprint.
Reporting observations rather than insights. "User 1 clicked the wrong tab" is a log, not a finding. "Three of five users navigated to the wrong section because 'Account' was interpreted as personal settings rather than company settings" is a finding — it names the cause, the element, and the severity.

Connected ideas

Usability testing is the primary evaluative method for qualitative behavioural data — complemented by quantitative methods for statistical questions and generative methods for discovery.

The most important pairing is Usability Testing with Prototyping. Testing without a prototype tests a finished product, producing rework. Prototyping without testing produces untested assumptions. Together they form the validate-before-you-build loop.

Run it right now

⏱ 10 minutes · Solo · No prep

The Corridor Test

1. Identify the single most important task a new user needs to complete. Write it as a task scenario in user goal language with no interface terminology.

2. Find one person who has never used your product. Give them a device with your product open and read them the task. Do not help. Do not explain. Watch.

3. Note the first moment they hesitate, go back, or take an unexpected path. Write down what they were trying to do, what they clicked instead, and what the interface showed.

4. Ask: if five more people encountered this same moment, would they all hesitate in the same place? If yes, you have identified a usability problem. If this is the first time you have watched someone outside the team use your product, this exercise has already produced more reliable data than any internal review.

10 minutes