Task: Stop receiving weekly summary emails ▸ Dashboard ▸ Projects ▸ Team ▸ Settings ▸ Notifications ✓ ▸ Reports first click: Team (wrong) corrected → Notifications ✓ tree-testing · strip the design · validate the structure · findability before the interface

Tree Testing

Validate your navigation structure before you design the interface around it. Tree testing strips away all visual design and asks users to find items in a text-only hierarchy — revealing whether the structure itself works.

Information ArchitectureNavigation ValidationTaxonomy TestingSite StructureSettings HierarchyHelp Centre Design

Two sentences.

Tree testing is a remote research method where participants navigate a text-only hierarchy — stripped of all visual design — to find specific items, with paths and success rates recorded. By removing all interface cues, it isolates the navigation structure as the variable, producing clean data on whether the IA supports findability independent of labels or layout.

The method became widely accessible through tools like Treejack (Optimal Workshop). A tree test can be designed, distributed to fifty participants, and analysed within a week — producing quantitative findability data that supports IA decisions with evidence rather than opinion.

Apply this when…

Card sorting has produced a proposed structure and the team needs to validate it under task conditions
An existing navigation is producing findability complaints or high search usage suggesting users cannot browse
A redesign needs confirmation that the new structure improves findability versus the current one
Two competing IA proposals exist and the team wants empirical data to decide
A help centre or documentation site is being restructured and needs taxonomy verification

When NOT to apply it

Skip it when the navigation has fewer than ten items (too simple), when the problem is label clarity not structure (different test needed), when you need to understand why users struggle (pair with think-aloud), or when the product has no hierarchical navigation.

The mechanism

Tree testing shows users only text labels, no visual interface. Participants rely entirely on understanding what each label means and how categories relate. Success and failure is attributable directly to structure and labelling, not visual design.

01
Findability is a structural property
Users navigate by category matching — scanning for the category they expect their target to be in. When the expected category is absent or mislabelled, they explore adjacent categories or abandon. Tree testing isolates this structural behaviour by removing everything else.
02
Task scenarios drive the test
Each task must describe a user goal without echoing navigation labels. "Find where to change notification settings" is poor if "Notifications" is a nav item. "You want to stop receiving weekly summary emails" is good — it describes the goal and lets the structure succeed or fail.
03
First click data is more predictive than overall success
A user who clicks the right category first succeeds 87% of the time; a user who clicks wrong first succeeds only 46%. First-click accuracy is the most actionable output — high first-click error rates indicate structural or labelling problems at the top level that propagate throughout the entire experience.
04
Success rate, directness, and first-click accuracy
Success rate = percentage who found the target. Directness = whether they found it without backtracking. First-click accuracy = which top-level category they tried first. A task with 40% success and low directness has a different problem than one with 40% success and high directness — the first is structural confusion, the second is label ambiguity at a lower level.

Tree testing validates, it does not design

The correct sequence: card sorting generates a structure grounded in user mental models, then tree testing validates it under task conditions. Teams that skip card sorting and go directly to tree testing often discover failure but have no user-grounded alternative to build from.

Atlassian Confluence's two-stage IA validation

When Atlassian restructured Confluence's navigation, the team first conducted card sorting to understand how users organised the product's concepts. The card sort revealed users grouped features around work contexts rather than product feature areas.

The proposed IA was then tree tested with task scenarios from real use cases. The tree test identified two category labels causing significant first-click errors — before any visual design. Fixing those labels at the IA stage cost an afternoon; fixing the same problem after a visual design had been built would have required a redesign sprint.

Atlassian Confluence · Navigation Restructure
Two-stage IA validation — card sort generates, tree test validates
Stage 1: Card SortgeneratesStage 2: Tree Test▸ My Work▸ Spaces ← first-click error▸ Templates▸ Admin ← label confusionTwo labels caught before visual designFix at IA stage: 1 afternoonFix after visual design: 1 sprintAtlassian Confluence · card sort → tree test · IA validated before design begins
Two-stage validation → problems caught at lowest cost

Test yourself & see real examples

No examples yet — be the first.

Spotted a navigation that was clearly never tested for findability — or one where the hierarchy feels so intuitive it must have been validated? Submit what you observed.

✓ Reviewed before publishing✓ Your name on every example you submit✓ Violation or fix — both welcome

Seen Tree Testing skipped in a real product? Help grow the evidence base.

Where teams go wrong

Writing tasks that echo navigation labels. "Navigate to Settings and find Notifications" is answered by word matching, not structure understanding. Use goal language: "You want to stop receiving weekly emails." If any task word appears as a nav label, rewrite it.
Testing a structure without prior card sorting. When a tree test fails, the team needs a user-grounded alternative. Without card sort data, the temptation is to adjust by intuition — which may produce another failed structure. Card sorting provides the mental model data that tree testing validates.
Confusing structural problems with labelling problems. A 40% success rate could mean the item is in the wrong category (structural) or the label does not communicate what it contains (labelling). Look at which categories participants visited first — related categories suggest labelling ambiguity; unrelated ones suggest structural mismatch.
Running tree tests without qualitative follow-up. Metrics show that something is wrong but not what participants were thinking. A short follow-up question or think-aloud with five participants provides the context that turns quantitative signal into a specific design direction.

Connected ideas

Tree testing occupies a specific, non-substitutable role in the IA validation toolkit.

The most important pairing is tree testing with card sorting. Card sorting without tree testing generates a structure that feels grounded but has never been tested under task conditions. Tree testing without card sorting validates a structure designed without user input. Run them sequentially: card sort to generate, tree test to validate.

Run it right now

⏱ 10 minutes · Solo · No prep

The Flat List Test

1. Open any product you use regularly. Write down every item in the top-level navigation as a flat text list, exactly as labelled.

2. Without navigating, predict: where would you find notification preferences? Billing information? Help documentation? Note which item you would click first for each.

3. Now actually navigate. For each task, note whether your predicted first click was correct and how many clicks to reach the destination.

4. Any task where your predicted first click was wrong — even as a regular user — is a tree testing signal. If you clicked the wrong category, new users almost certainly do too. A tree test with real users would confirm and quantify the problem.

10 minutes