
Heuristic Evaluation
✓ a ranked list of problems in an afternoon
On this page
- TL;DR
- What is heuristic evaluation?
- The 10 usability heuristics
- When to use heuristic evaluation
- When it is the wrong method
- How to run a heuristic evaluation, step by step
- How to write findings that get fixed
- Common mistakes
- What to do with the findings
- Heuristic evaluation vs usability testing
- Templates and tools
- The bottom line
Most design teams already know where their product hurts. They just cannot prove it, and they cannot get it prioritized.
Heuristic evaluation is the cheapest way to turn that shared hunch into a written, ranked list. Two or three people who know usability principles walk the interface, note every place it breaks one, and rate how badly. No recruitment, no scheduling, no incentives. You can run one in an afternoon.
It also has a real limit, and the limit matters more than most articles admit. Heuristic evaluation tells you what an expert thinks is wrong. It does not tell you what actually costs your customers anything. This guide covers both halves: how to run one properly, and what to do with the list once you have it.
TL;DR
Heuristic evaluation is a structured expert review. A small group of evaluators inspects an interface against an agreed set of usability principles, logs every violation, and rates each one for severity. It is fast and cheap because it needs no participants. It is best used early, to clear out obvious problems before you spend money watching real people hit them.
To run one:
- Agree the heuristics and the scope
- Brief 3 to 5 evaluators
- Have each person review alone first
- Combine the findings into one list
- Rate severity together
- Hand over a ranked list, not a document
What is heuristic evaluation?
Heuristic evaluation is a usability inspection method. A small number of evaluators examine an interface and judge it against a set of established usability principles, called heuristics. Anywhere the interface breaks one, they record it as a finding.
The method was developed by Jakob Nielsen and Rolf Molich in 1990, and the 10 heuristics published by Nielsen Norman Group remain the default framework almost everyone uses.
The word heuristic is doing real work. These are rules of thumb, not laws. They are broad enough to apply to a checkout flow, a dashboard and a mobile onboarding sequence, which is what makes them useful, and vague enough that two competent evaluators will disagree, which is why you never run one alone.
What separates it from someone simply having opinions about the design is structure. Every finding attaches to a named principle, sits at a specific location in the product, and carries a severity rating. That is what makes the output arguable rather than personal.
The 10 usability heuristics
Paraphrased from Nielsen Norman Group's framework. The originals are worth reading in full.
- Visibility of system status. The product tells people what is happening. Loading states, progress indicators, confirmation that a thing saved.
- Match between the system and the real world. The language is the user's, not the database's. Concepts appear in an order that makes sense outside the org chart.
- User control and freedom. People pick wrong options. There is a clearly marked way out, and undo exists.
- Consistency and standards. The same action looks the same everywhere, and platform conventions are respected rather than reinvented.
- Error prevention. Better to make the mistake impossible than to apologize for it afterwards. Constrain the input, confirm the destructive action.
- Recognition rather than recall. People should not have to remember something from a previous screen. Show the options.
- Flexibility and efficiency of use. Shortcuts for people who are here every day, without cluttering the path for people who arrived this morning.
- Aesthetic and minimalist design. Every extra element competes with the one that matters. Cut what is merely nice to have.
- Help users recognize, diagnose and recover from errors. Plain language, no codes, and a stated next step.
- Help and documentation. Ideally unnecessary. Where it is necessary, it is findable and specific to the task in hand.
Some teams extend the list for their context, adding accessibility or trust and privacy heuristics. That is fine, as long as everyone evaluates against the same list. A shared list is the whole point.
When to use heuristic evaluation
It earns its place in four situations.
- Before you spend money on participants. Watching five people trip over a broken back button is an expensive way to learn that your back button is broken. Clear the obvious violations first, then put real people in front of what remains.
- When you have inherited something. New role, new product, or an acquisition. A structured pass gives you a map of the problems and a defensible starting position.
- On competitor products. You can evaluate anything you can access, which makes this one of the few methods that works on software you do not own.
- When there is no budget and no time. Heuristic evaluation is the method you can always run. Three colleagues and two hours is a real evaluation.
Where it fits in the wider set of UX research methods is early and cheap. It narrows the field. It does not settle anything.
When it is the wrong method
Heuristic evaluation produces expert opinion. Treat it as evidence and it will mislead you in three predictable ways.
It cannot tell you what people actually do
An evaluator knows the product too well to be surprised by it. The confusion a first-time user feels at step three is invisible to someone who built step three. If the question is behavioral, you need usability testing.
It cannot tell you what people want
Heuristics measure an interface against principles of good interface design. They say nothing about whether the underlying feature is worth having. That question needs user interviews.
It over-reports
Evaluators find violations that no real person would ever notice, and severity ratings drift upward because everything looks important when you are staring at it. Expect a long list and expect to cut it.
The reliable pattern is sequencing. Use heuristic evaluation to generate hypotheses about where the product is failing. Use research with real people to find out which of those hypotheses cost you anything. Great Question is for the second half: recruiting your own customers, running the sessions, and keeping what you learn somewhere the next person can find it.
How to run a heuristic evaluation, step by step
1. Define the scope
Pick a flow, not a product. “Checkout, from cart to confirmation” is a scope. “The app” is a wish. A tight scope is what makes findings comparable between evaluators.
Write down the user goal, the entry point, the device and the account state. Evaluating a signed-in enterprise account and a fresh trial account produces two different sets of findings.
2. Choose the heuristics
Nielsen's 10 by default. Add to them only if your context genuinely demands it, and share the final list before anyone starts. Evaluators working from different lists produce findings you cannot merge.
3. Recruit 3 to 5 evaluators
Nielsen's own research on this is the reason the method is structured the way it is: individual evaluators miss most problems, and the returns from adding people flatten out after about five. Three is the working minimum. One is not an evaluation.
They do not all need to be usability specialists. A mix of a designer, a researcher and someone with deep domain knowledge tends to surface a wider range than three specialists would.
4. Brief everyone properly
Give each evaluator the scope, the heuristic list, the account credentials, the severity scale and the logging template. Then say the part people forget: review alone, and do not discuss findings until everyone has finished.
Independence is the mechanism. The moment two evaluators compare notes mid-review, you have one opinion with two names on it.
5. Review in two passes
First pass, move through the flow at normal speed to get a feel for it. Second pass, go slowly and inspect each screen against every heuristic in turn.
The two-pass structure exists because the first walkthrough tells you how the product feels and the second tells you why. Doing only the slow pass produces a list of technically correct nitpicks with no sense of what matters.
For each finding, log where it is, which heuristic it breaks, what the user would experience, and a screenshot.
6. Merge the findings
Bring everything into one list. Deduplicate, but keep a count of how many evaluators independently flagged each issue. That count is a useful signal in its own right, and it is the number that tends to persuade stakeholders.
7. Rate severity together
Merging is clerical. Severity is a judgment call, and it should be made as a group, after everyone has seen the full combined list. Rating in isolation produces wildly inconsistent scores because nobody has a sense of the range yet.
A standard four-point scale, adapted from Nielsen's:
- 0, not a problem. Disagreement between evaluators, or a violation with no practical consequence.
- 1, cosmetic. Fix if there is spare capacity.
- 2, minor. Users are mildly slowed. Low priority.
- 3, major. Users struggle. Fix in the next cycle.
- 4, catastrophic. Users cannot complete the task. Fix before ship.
Rate on frequency, impact and persistence. A small annoyance on a screen every customer sees every day can outrank a serious problem in a corner nobody visits.
8. Hand over a ranked list
Not a report. A ranked, filterable list of findings with severity, location, heuristic and screenshot, in whatever tool your team actually plans work in.
The evaluation is only worth the time if the findings get fixed, and a 40-page PDF is where findings go to be admired rather than actioned.
How to write findings that get fixed
Most heuristic evaluations fail here rather than in the review itself.
- Name the principle, not your preference. “Breaks consistency: this button is labeled Save here and Update on the other three screens” travels further than “this feels off.”
- Describe the user consequence. Not “no loading state” but “no loading state, so people submit twice and get duplicate records.”
- Separate the finding from the fix. Log what is broken. Suggest a fix if you have one, clearly marked as a suggestion. Merging them invites an argument about your solution instead of a conversation about the problem.
- Keep the disagreements. If two evaluators split on something, record it as contested rather than dropping it. Contested findings are often the best candidates for testing with real people, because they are exactly the ones expert judgment cannot settle.
Common mistakes
- Running it alone. One evaluator finds a fraction of the problems and has no way to know which fraction.
- Reviewing as a group. Kills independence. Group discussion belongs in the merge and severity stages, not the review.
- Scoring severity while you review. Ratings drift because there is no context yet. Do it once, at the end, together.
- Treating the output as evidence. It is a hypothesis list. Say so when you present it, or someone will quote it back to you as user research.
- Evaluating against no agreed heuristics. Without a shared list you have collected opinions and given them a methodology-sounding name.
- Stopping there. The evaluation identifies candidates. Confirming which ones matter is a separate job, and it needs real people.
What to do with the findings
Split the ranked list three ways.
Fix immediately
Severity 3 and 4 with an obvious remedy. Broken back buttons, missing error messages, destructive actions with no confirmation. These need no further research.
Test with real people
Contested findings, anything where the fix is not obvious, and anything where evaluators disagreed on severity. Put a prototype in front of your own customers or run unmoderated tests on the current build and find out whether the problem is real.
Park
Cosmetic findings, and anything in a flow nobody uses. Keep the list. Revisit when that area comes up for work.
That middle group is where the method pays off. You have gone from a vague sense that onboarding is bad to a specific, ranked set of questions worth spending participant time on. Great Question handles that half: recruiting from your own customer base, scheduling, incentives, and a repository where the findings sit alongside everything else you know, so the next person to evaluate this flow starts from what you already learned rather than from scratch.
Heuristic evaluation vs usability testing
The two get compared constantly, usually as though you should pick one.
Heuristic evaluation | Usability testing | |
|---|---|---|
| Who takes part | 3 to 5 evaluators | Real users, typically 5 or more |
| What it produces | Expert opinion on interface quality | Observed behavior |
| Time to run | Hours to days | Days to weeks |
| Cost | Internal time only | Recruitment, incentives, scheduling |
| Best at | Finding candidate problems fast | Confirming which problems are real |
| Blind spot | Cannot see genuine user confusion | Cannot cover every screen |
They answer different questions. Heuristic evaluation gives you breadth cheaply: an expert can inspect every screen in a flow in an afternoon. Usability testing gives you truth expensively: five people can only attempt a handful of tasks, but what they do is real.
Run the evaluation first. Let it decide what the testing should focus on.
Templates and tools
You need less than you think. A shared spreadsheet with columns for location, heuristic, description, severity, evaluator and screenshot covers the whole method. Great Question's template library has starting points for the research that follows the evaluation.
For the follow-up work, the platform handles recruiting your own customers, screening, scheduling and incentives, plus prototype tests, card sorting and unmoderated studies. Findings land in the repository with the supporting clips attached, so a claim in a readout can be traced back to the moment someone actually struggled.
Note that no tool runs the evaluation for you. It is a human judgment method, and automated accessibility or performance scanners are a different thing entirely. Useful, but not this.
The bottom line
Heuristic evaluation is the highest-return two hours in design research, as long as you are honest about what you are buying. You get a fast, structured, ranked list of everywhere your interface breaks established principles, produced by people who know what good looks like.
What you do not get is proof. The list is a set of hypotheses about where users struggle, generated by people who are not users. Fix the obvious violations, test the contested ones with your own customers, and the two methods together will get you further than either does alone.
Frequently asked questions
What is heuristic evaluation in UX?
Heuristic evaluation is a usability inspection method where a small group of evaluators reviews an interface against a set of established usability principles, records every violation, and rates each one for severity. It needs no research participants, which makes it one of the fastest and cheapest methods available.
How many evaluators do you need for a heuristic evaluation?
Three to five. A single evaluator misses most of the problems in an interface, and the number of new problems found drops off sharply after about five people. Three is the practical minimum for a result you can rely on.
What are the 10 usability heuristics?
Visibility of system status; match between system and real world; user control and freedom; consistency and standards; error prevention; recognition rather than recall; flexibility and efficiency of use; aesthetic and minimalist design; helping users recognize, diagnose and recover from errors; and help and documentation. The framework comes from Nielsen Norman Group.
What is the difference between a heuristic evaluation and a UX audit?
A UX audit is the wider program: scoping, choosing what to assess, running the work and producing a remediation plan. Heuristic evaluation is one method used inside it. An audit might combine heuristic evaluation with analytics review, accessibility checks and user testing.
Is heuristic evaluation qualitative or quantitative?
Qualitative. It produces descriptive findings based on expert judgment. Severity ratings add a numeric layer for prioritization, but they are structured opinion rather than measurement.
Can heuristic evaluation replace user testing?
No. It identifies where an interface breaks usability principles, which is not the same as knowing where real people struggle. The two are complementary: evaluate first to narrow the field, then test with real people to find out which problems actually matter.




