Tree testing: a practical guide

What tree testing measures, how to design tasks and a tree, which metrics to read, and how to run a tree test in Versive.

Updated
8 min read

Tree testing measures whether people can navigate using labels and structure alone, without visual design as a cue. Use it to evaluate a new site map, an app's settings menu, or a help center's category structure. The method requires realistic tasks, a representative tree, and analysis of both success and the path each participant took.

What tree testing measures

A tree test strips a navigation structure down to text: a list of nested categories and items, with no icons, no visual hierarchy, no styling to guide the eye. You give a participant a task, something like "find where you'd change your billing plan," and watch which path they take through the hierarchy to answer it.

Because participants have only labels and structure, a tree test isolates the information architecture. If they cannot find the correct node, visual details such as button size or color are not responsible. The item's category or label may not match how participants think about the task.

That makes tree testing well suited to a specific moment in a project: after you have a candidate structure but before you've invested in visual design around it, so you're not paying to redo layouts because the underlying navigation was wrong to begin with.

Tree testing vs. card sorting

Tree testing and card sorting get grouped together often, and they do pair well, but they answer different questions at different points in the process.

Card sorting is generative. You hand participants a set of cards, usually the individual pages or features that need to live somewhere, and ask them to group those cards into categories that make sense to them (in an open sort, they can also invent their own category names). The output is raw material you use to draft a structure.

Tree testing is evaluative. You've already built a structure, and now you're checking whether it works by asking people to navigate it to complete specific tasks. Participants only move through the hierarchy you built, picking a node at each level until they land on an answer or give up.

Card sortingTree testing
PurposeGenerate a structureEvaluate a structure
What participants doGroup cards into categoriesNavigate a tree to complete tasks
Best timingBefore you've drafted an IAAfter you've drafted an IA, before visual design
What it revealsHow people naturally categorize your contentWhether people can find specific things in what you built

The two methods often run in sequence: a card sort generates candidate structures, then a tree test evaluates the selected candidate before visual design begins. Card sorting: a practical guide explains the generative part of that sequence.

Designing tree test tasks

Task wording strongly affects a tree test, so review it carefully before building the tree.

A good task describes a goal in the participant's own terms, not your navigation's terms. "Find where you'd change your billing plan" works because it describes something a person wants to do. "Find the Billing section" doesn't, because it hands them the answer by naming the label you're trying to test.

A few practical guidelines:

Order matters too: later tasks benefit from a familiarity with the tree that the first task didn't have. Randomizing task order across participants spreads that learning effect out instead of concentrating it on whichever task happens to be last.

Building the tree

The tree itself should mirror the real structure you're testing, or a serious candidate for it, not a simplified stand-in. Leaving out sections to keep the tree tidy also removes the plausible wrong turns participants would take in the real product, which quietly inflates your success rates.

A few things to get right when building it:

Reading the results: success rate and directness

Two metrics do most of the interpretive work in a tree test, and they answer different questions.

Success rate is the simplest read: what share of participants landed on a correct answer for a given task, versus a wrong answer or giving up entirely. It tells you whether a task, as a whole, works. A task with a low success rate needs attention somewhere in its path.

Directness tells you where in that path things went wrong. A participant who goes straight to the correct node without doubling back is direct; one who picks a category, backtracks, tries another, and eventually lands on the right (or wrong) answer is not. Low success with high directness usually means participants made a fast, confident, wrong choice, which points to a mislabeled or miscategorized node. Low success with low directness points to a structure that's confusing throughout the path, not just at one decision point.

Read the two metrics together rather than treating every failed task alike. If participants confidently choose the same wrong category, revise the label or parent category. If they wander and give up, the structure may be too deep or ambiguous across several levels.

It's also worth looking at the actual paths participants took, not just whether they succeeded. A handful of participants converging on the same wrong node is a much stronger signal than an even scatter of different wrong answers, even if the success rate is identical in both cases.

Running a tree test in Versive

Versive includes Tree Test as a question type, so a tree test can run as its own study or sit inside a larger session alongside surveys, ratings, or AI interview questions.

To build the tree, work in visual mode, dragging and dropping nodes to construct and rearrange the hierarchy directly, or in text mode, typing or pasting the tree as indented text. Text mode is the faster path if you're starting from an existing site map or spreadsheet. Trees can also be imported and reused across studies, so a structure you validate once doesn't need to be rebuilt for the next round of testing.

Each task gets a description and one or more correct answers, matching the task design guidance above: write the description around a real goal, and mark every node that would count as a legitimate answer. You can randomize task order across participants, and configure whether they're allowed to skip a task, go back to a previous step, or select a parent node as their final answer rather than drilling all the way down. Tree depth defaults to 6 visible levels, which covers most real navigation structures without extra configuration.

Results land in a dedicated dashboard: Sankey diagrams showing how participants moved through the tree, including where they backtracked; path analysis breaking down the most common routes per task; success rates split into correct, incorrect, and gave-up; and per-task analytics showing which parts of the structure need rebuilding. Success rates come straight out of the dashboard, and the Sankey diagrams and path analysis are where you read the backtracking behavior that separates a direct path from a wandering one.

Because Tree Test is a question type like any other, it composes with the rest of a study: open with a card sort to generate structure candidates, follow with a tree test on the resulting hierarchy, and close with an AI Question asking participants why they picked the path they did, all in one session. See the question types reference for how tree test settings sit alongside every other type in the builder.

When to run a tree test

Run a tree test after drafting a structure and before building its visual design. At that stage, you can correct labeling and categorization problems before they affect layouts or ship in a redesign.

Pair it with a card sort when you need to generate the structure. Once visual design is in place, follow with a broader usability test to confirm that the structure works with real content. User interviews vs. surveys compares structured and conversational methods.

Add a Tree Test question to a study, build or import the hierarchy, and start with two or three tasks for the sections you are least confident about.

Frequently asked questions

What is tree testing used for?

Tree testing checks whether people can find things in your navigation structure, using a text-only version of the hierarchy with no visual design to help or hide problems.

What is the difference between tree testing and card sorting?

Card sorting is generative: participants group items to help you build a structure. Tree testing is evaluative: participants navigate a structure you already built to see if it works.

What is a good success rate for a tree test task?

There is no universal number, since it depends on the task and the stakes of getting it wrong, but a task where most participants fail or give up is a strong signal that node or label needs to change.

What is directness in a tree test?

Directness measures whether a participant went straight to the correct node or backtracked along the way. A low success rate with high directness usually means a clear but wrong choice, while backtracking points to a confusing structure.

Can I combine tree testing with other question types in the same study?

Yes. A tree test is a question type in Versive, so you can place it alongside surveys, AI interview questions, or other tasks in a single study.

Full reference

Question types


Keep reading

Faster research, better insights. Start now.