
Findability testing
✓ check users can navigate your IA before you build it
On this page
- TL;DR
- What is tree testing?
- Common tree testing tasks
- Benefits of tree testing
- Limitations of tree testing
- When to use tree testing
- Common tree testing use cases and examples
- Qualitative vs quantitative tree testing
- How many participants do you need?
- How to run a tree test
- How to analyze tree test data
- Best practices for tree testing
- Tree testing vs card sorting vs first-click testing
- The best tree testing tools in 2026
- The bottom line
TL;DR
Tree testing, also called reverse card sorting, checks whether people can find things in your site or app using the navigation labels alone, with no visual design to lean on. You give participants a text-only version of your navigation, set a few realistic tasks, and watch where they click. The numbers that matter are success rate, directness, and time on task. Run it when you are planning a redesign, validating a card sort, or adding a new section to your information architecture.
Your visitors rely on your information architecture, or IA, to get anything done on your site. When the way content is organized and labeled does not match how they actually think, they hesitate, take wrong turns, or give up altogether. You lose the signup or the sale, and you rarely find out why.
Tree testing is how you catch those problems before a new structure ships. This guide covers what tree testing is, when to run it, the metrics to watch, realistic sample sizes, and the tools worth using in 2026.
What is tree testing?
Tree testing, also known as reverse card sorting, is a UX research method used to evaluate the findability of a website or app's information architecture. You strip away the visual design and present participants with only the navigation structure, or tree, then ask them to work through it to complete specific tasks. Because the design is gone, you see whether people can actually find things through the intended paths, not whether visual cues helped them get there.
The point of a tree test is to find the weak spots in your IA: categories that do not sit where people expect, labels that are unclear, and terminology that makes sense internally but not to your users.
The paths participants take, how often they succeed, and how long each task takes all point to where your structure is confusing. That gives you specific, prioritized fixes instead of a vague sense that navigation could be better.
Common tree testing tasks
A tree test is only as good as its tasks. Each one asks a participant to find something specific, testing whether the labels and hierarchy lead them to the right place. Common task types include:
- Locating specific information. Ask participants to find a particular page or piece of content. Example: “Find the company's return policy.”
- Category selection. Ask where they would expect something to live. Example: “Where would you look for battery replacement information for a product?”
- Choosing between plausible paths. Give participants a task where more than one part of the navigation might initially seem right. Their path shows whether your hierarchy makes the intended destination clear. Example: “Find where you would go to update your product's software.”
- Testing terminology. Use tasks that reveal whether people understand the labels in your navigation without explaining those labels in the task itself. For example, if your navigation uses “Corporate Social Responsibility,” ask participants to find information about the company's environmental or community initiatives rather than telling them to look for “CSR.”
Benefits of tree testing
Find navigation problems before they become usability problems
Tree testing shows you where users hesitate, take a wrong turn, or hit a dead end in your navigation, without the visual design masking the problem. You get a clear picture of which parts of the hierarchy need work so people can find what they came for faster.
Evaluate your language and labels
Labels carry a lot of weight, and tree testing examines them in isolation, free from imagery or interactive cues. Jargon, ambiguity, and phrasing that makes sense internally but not to customers become easier to spot. You see whether your labels work without design doing the explaining for them.
Understanding user mental models
As people move through a bare tree, you learn where they expect information to live and which categories they associate with it. That tells you the pathways they expect, so you can shape your structure around their thinking rather than your org chart.
Ranking usability issues by impact
Not every problem deserves the same attention. Some labels trip up nearly everyone; others only affect a handful of people. Because tree testing gives you success rates, directness, time on task, and navigation paths, you can see which problems are widespread and prioritize accordingly.
Limitations of tree testing
Like any method, tree testing has limits worth knowing.
No visual context
Tree testing evaluates the IA on its own, with no design or interactive elements. That keeps the assessment focused, but it does not tell you how imagery, layout, or visual hierarchy will affect navigation once the real interface is in front of users.
It doesn't capture the full user experience
Tree testing focuses on whether people can locate information within a hierarchy. It doesn't tell you how they interact with the finished interface, understand the content they find, or respond to the overall experience. Pair it with other methods when you need that broader picture.
It assumes a specific task
Real users show up with their own goals and motivations. A tree test assigns tasks, which can miss the messier, more varied intent people bring to a live product. Your results tell you how well the structure supports the tasks you tested, not every possible way someone might navigate it.
When to use tree testing
Tree testing earns its place whenever you need to check that your labels and structure hold up, whether you are building something new or refining what you have.
When starting a redesign
Early in a redesign, a tree test on your current IA shows where the existing structure already fails. While broader usability tests tell you about visual design and interaction, tree testing zeroes in on the IA and gives you a baseline to compare against as the structure evolves.
After a card sorting exercise
Card sorting surfaces how people group content, but it stops short of handing you a finished structure. Tree testing is the natural follow-up: it validates the hierarchy a card sort suggested and tells you whether people can actually navigate it.
Before content and layout work
Testing the hierarchy and labels before anyone writes copy or builds screens means you can validate the foundation while changes are still relatively easy to make. That means fewer surprises later and fewer expensive reworks once design and development are underway.
After major updates or new features
When you add a section or feature, a tree test checks that people can find it and that it sits sensibly inside the existing structure, rather than becoming something users struggle to discover.
Common tree testing use cases and examples
Validating a new IA design
Before rolling out a new structure, test it. An ecommerce site reworking its product categories can run a tree test to confirm people can find the products they need under the new labels.
Comparing multiple IA proposals
When several structures are on the table, tree testing lets you compare them based on how people actually navigate them. A news site weighing different ways to organize articles can test each and see which one readers navigate most easily.
Identifying problematic labels
If you suspect certain terms are unclear, a tree test can show you where people get stuck. A health portal can check whether labels like “cardiovascular” or “hematology” work for a general audience or need plainer wording.
Testing new features or sections
A social platform adding a “marketplace” section can tree-test the updated navigation to confirm people find it naturally, rather than assuming the new label is clear enough on its own.
Qualitative vs quantitative tree testing
You can run a tree test qualitatively or quantitatively, and the two answer different questions.
Qualitative tree testing
Qualitative tree testing digs into the why. You watch a small number of people navigate and ask follow-up questions when they hesitate or backtrack. Techniques like think-aloud protocols surface the reasoning behind a wrong turn, which is often where the most useful insight lives. A smaller sample works here because the goal is to uncover problems and understand why they happen, not produce statistically reliable success rates.
Quantitative tree testing
Quantitative tree testing is about patterns across a larger group. You track success rates, time on task, directness, and the paths people take, then look for patterns across participants. This is where sample size starts to matter, which we cover below.
How many participants do you need?
This is where a lot of guides oversimplify. The honest answer is that it depends on what you need the results to tell you.
For qualitative tree tests, you can start with a small group and add participants until the major patterns become clear. The goal is to uncover obvious breakages and understand the reasoning behind them, not to calculate statistically reliable percentages.
For quantitative tree tests, you need a larger sample so your success and directness numbers are more stable. Around 50 participants can be a useful starting point for a directional quantitative study, but there is no universal number that works for every tree test. The sample you need depends on the size of your audience, the number of segments you are comparing, and how much confidence you need in the results.
A practical middle path: run a small qualitative round first to catch confusing task wording and obvious label problems, then launch the larger quantitative study once the test itself is clean.
How to run a tree test
Define your objectives
Decide what you are trying to learn before anything else. Validating a new IA, finding navigation problems, or comparing two structures will each produce a different study, so be specific.
Draft your tree
Build a text-only version of your existing or proposed IA. No visuals, no styling, just the labels and categories you want to test.
Design the tasks
Write realistic tasks that mirror what people actually do with your product. Keep them clear and free of leading language. For example: “Find where you would get pricing information for enterprise teams.” Keep the number manageable so participants do not fatigue and start clicking without thinking. Around 10 tasks is often plenty, though the right number depends on their complexity.
Recruit your participants
Your participants should reflect your real audience. If you are testing an existing product, recruiting from your own customers gives you participants with firsthand context, while screener questions help you reach the right subset. A relevant research incentive can improve participation, and if you skip incentives, plan to invite a wider pool to hit your numbers.
Recruitment can also become a bottleneck as research volume grows. At ServiceNow, moving research to Great Question and recruiting from its own customer base helped cut participant recruitment from 118 days to 6.
Conduct the test
Using a tree testing tool, invite participants to work through the tree and complete the tasks. Capture where they click, how long they take, and where they backtrack. Quantitative tree tests are well suited to unmoderated research, so participants can complete them on their own time while you collect results at scale.
Gather feedback
Beyond the click paths, ask people what tripped them up. A quick post-task question about confusing labels can explain why participants took a particular path or failed a task.
Analyze the data
Look at success rate, directness, time on task, and the paths people took, then identify the labels, categories, and branches where participants struggled or took unexpected routes.
How to analyze tree test data
Once the test wraps, your results sit in your tree testing software, ready to read. A few metrics do most of the work.
Success rate is the share of people who completed a task correctly. Rather than aiming for a universal benchmark, compare success across tasks and look closely at the ones where participants consistently struggle. A lower success rate on one task than the rest is a strong signal that something in that part of the structure needs attention.
Directness is the share of people who reached the answer without backtracking. Low directness can point to ambiguous labels or competing paths, even when participants eventually reach the right place.
Time on task is how long each task took. Spikes can point to hesitation and confusion, especially when they line up with low success or directness.
Path analysis traces the routes people took, which exposes the categories that pull people in the wrong direction.
Do not stop at the numbers. Qualitative feedback tells you which label confused someone and why, which is what turns a low score into an actual fix.
Pulling patterns out of dozens of open-ended responses is the part that usually eats an afternoon. In Great Question, AI theme-clustering groups that feedback automatically, and a unified results table puts the quantitative metrics and qualitative themes in one view, so you are reading one screen instead of stitching together a spreadsheet and a doc.
Best practices for tree testing
Recruit the right people
Match your participants to your real audience. Relevant participants give you findings you can actually use to make decisions about the structure.
Write realistic tasks
Base tasks on things people genuinely do with your product. The more real the task, the more useful the behavior you observe.
Avoid leading questions
Do not write tasks that hint at the answer. A task that names the target category makes it easier for participants to guess where to go and undermines the test.
Keep the task count sensible
Keep the study focused enough that participants can give each task proper attention. Around 10 tasks is often enough to cover the important parts of a tree without making the session drag.
Pilot before you launch
Run a small pilot first, ideally with a few participants, to catch confusing wording or problems in the tree itself before you commit your full sample. It is much easier to rewrite a broken task before the full study than after dozens of people have completed it.
Combine with other methods
Tree testing is strongest alongside card sorting or first-click testing. Together they give you a fuller read on structure, labels, and how people navigate toward what they need.
Tree testing vs card sorting vs first-click testing
These three methods are often mentioned together because they all help evaluate navigation and information architecture, but they answer different questions.
Tree testing evaluates findability: can people navigate your existing or proposed structure to locate specific items?
Card sorting helps you understand how people group and categorize information. Participants organize content in ways that make sense to them, which reveals their mental models and can help you build or revise the IA.
First-click testing focuses on the first step someone takes toward completing a task. By analyzing where people click first, you learn whether the interface is sending them in the right direction from the start.
Used together, the three give you a broader view of how people understand, navigate, and find information in your product.
The best tree testing tools in 2026
There are several solid tools for running tree tests, and the right choice depends on whether you need a dedicated point solution or something that connects the test to the rest of your research.
Optimal Workshop is one of the best-known options for dedicated information architecture research. Its Treejack tool is built specifically for tree testing, with detailed path analysis and visualizations that help you see where participants succeed, backtrack, or head in the wrong direction. If tree testing and card sorting make up a large share of your research, that specialization can be useful.
UXtweak and Lyssna also support tree testing alongside related methods. UXtweak offers detailed path and first-click analysis, while Lyssna is geared toward quick, lightweight remote studies such as tree tests, first-click tests, and preference tests. Useberry is another option for teams that want tree testing alongside prototype and usability testing.
The trade-off with dedicated tools is what happens around the tree test. If you also need to recruit your own customers, screen participants, manage incentives, run follow-up interviews, and keep findings alongside the rest of your research, you may still need additional tools or workflows to connect those pieces.
Where Great Question fits
Great Question takes a broader approach. Tree testing runs as an unmoderated study inside the same platform you use to recruit participants, manage screeners and incentives, conduct interviews and surveys, analyze findings, and store research in a shared repository.
That matters because a tree test rarely answers every question on its own. A low success or directness rate can tell you where people struggled, but you may still need qualitative research to understand why. In Great Question, you can follow the tree test with a moderated or AI-moderated interview without moving the research into another system. The transcripts and findings flow back into the same research repository.
AI theme-clustering also helps connect the quantitative tree-testing results with participants' open-ended feedback. Instead of reviewing metrics in one tool and manually sorting comments somewhere else, you can see recurring themes alongside the results and trace findings back to the underlying evidence.
For teams that only need an occasional standalone tree test, a dedicated tool may be enough. But if tree testing is one method inside a larger research program, Great Question keeps recruitment, execution, follow-up, analysis, and the repository connected from the start.
The bottom line
Getting your information architecture right is foundational to a usable product, and tree testing is how you check your work before design and development make changes more expensive. Draft a clean tree, write realistic tasks, recruit people who reflect your real users, and read success rate, directness, and time on task together with what people tell you.
The teams that get the most out of it treat tree testing as one connected step in their research rather than a one-off in a standalone tool. Test the structure, follow up on what surprises you, and keep the results somewhere the rest of your team can actually use them.
Frequently asked questions
Is tree testing qualitative or quantitative?
It can be either. Qualitative tree tests use a small group and think-aloud follow-ups to understand why people navigate the way they do. Quantitative tree tests use a larger sample to measure patterns in success, directness, time on task, and navigation paths. Many teams run a small qualitative round first, then scale up.
How many tasks should a tree test have?
There is no single required number, but keep the study focused enough to avoid participant fatigue. Around 10 tasks is a useful starting point for many studies, with fewer if the tasks or tree are particularly complex.
What is a good success rate for a tree test?
There is no universal success-rate benchmark that works for every tree test. Look at task difficulty, compare performance across tasks, and pay particular attention to places where success drops or participants repeatedly take indirect routes. If you are comparing an old and new IA, improvement against your existing baseline is often more useful than an arbitrary cutoff.
What is the difference between tree testing and card sorting?
Card sorting helps you understand how people expect information to be grouped, which can inform an IA. Tree testing checks whether an existing or proposed IA actually works by measuring whether people can find things in it. They are complementary methods: card sorting can inform the structure, and tree testing can validate it.




