Short answer: choose the method that matches your research question, not your calendar. Remote unmoderated testing is fastest for validating known flows at scale. Remote moderated testing is the default workhorse for most product questions. Lab testing earns its cost when you need controlled conditions, hardware, or sensitive contexts. Contextual inquiry is for understanding the environment a product lives in. Guerrilla testing is for cheap, early gut-checks: nothing more, nothing less.
At Things, we run all four. "Research before opinion" is not a slogan we hang on the wall; it is how 50+ shipped products avoided expensive redesigns. Here is the decision framework we actually use.
What are the main usability testing methods?
There are four core methods: remote (moderated or unmoderated), lab-based, contextual, and guerrilla testing. They differ along three axes: control, context, and cost.
- Remote moderated: a facilitator guides a participant through tasks over a video call, screen shared, thinking aloud.
- Remote unmoderated: participants complete tasks alone via a testing platform; you review recordings and metrics afterward.
- Lab testing: participants come to a controlled environment; you observe directly, often with eye tracking, prototypes, or physical devices.
- Contextual inquiry: you observe people using the product (or doing the job your product will do) in their real environment: their office, warehouse, kitchen, car.
- Guerrilla testing: you intercept people in cafés, co-working spaces, or hallways and run five-minute sessions on the spot.
When should you use remote moderated usability testing?
Use remote moderated testing as your default for evaluating flows, prototypes, and new features, especially when you need to ask "why?" in the moment.
A moderator can probe hesitation, ask follow-up questions, and rescue a session when a prototype breaks.
Strengths: deep insight per session, flexible mid-session, wide recruitment reach, works with rough prototypes.
Weaknesses: scheduling overhead, moderator skill matters (leading questions poison data), participants are outside their natural context, tech issues eat session time.
Cost/speed profile: moderate cost, moderate speed. A five-session round can run inside a week including synthesis.
Choose it when: you are testing a redesigned checkout, a new onboarding flow, or an unfamiliar interaction pattern and need to understand reasoning, not just success rates.
When is remote unmoderated testing the right call?
Use unmoderated testing when your tasks are well-defined, your prototype is stable, and you need volume or speed more than depth.
Participants work alone, so there is no one to clarify a confusing task, which means your test plan must be airtight. Pilot it first.
Strengths: fast (results in hours, not weeks), scalable to dozens of participants, no scheduling, natural pacing without a moderator watching, good for benchmarking metrics like task success and time-on-task across releases.
Weaknesses: no follow-up questions, participants may satisfice or misread tasks, garbage-in-garbage-out on task wording, panel participants can be "professional testers."
Cost/speed profile: low-to-moderate cost, fastest turnaround of the four.
Choose it when: comparing two navigation labels, benchmarking a live product quarterly, or validating a flow you already understand qualitatively.
When does lab usability testing justify its cost?
Bring people into a lab when you need control, specialized equipment, or observation that remote setups cannot deliver.
Lab testing is the most expensive method per participant (facility, recruitment, incentives, travel, staffing), so it should answer questions the cheaper methods cannot.
Strengths: full environmental control, eye tracking and biometrics, testing physical products and multi-device setups, stakeholders can observe live behind the glass (nothing converts a skeptical executive like watching a real user struggle), sensitive data stays in the room.
Weaknesses: expensive, slow to organize, geographically constrained recruitment, and an artificial setting. People behave differently in a lab than on a crowded commuter train.
Cost/speed profile: high cost, slow. Budget two to four weeks per round.
Choose it when: testing hardware or kiosk interfaces, running eye-tracking studies, working under strict confidentiality, or testing with populations that need in-person support.
Contextual inquiry vs usability test: what's the difference?
A usability test evaluates whether people can use your design; contextual inquiry investigates how work actually happens in the real world. One is evaluative, the other generative.
In contextual inquiry, you go to the participant (a nurse's station, a logistics depot, a trader's desk) and observe them doing their real work with their real tools, interruptions included. You are not testing tasks; you are learning what the tasks are. The classic posture is master–apprentice: they work, you watch and ask.
Strengths: uncovers workarounds, environmental constraints, and needs users would never articulate in a lab; the single best input for defining what to build.
Weaknesses: time-intensive, small samples, requires access to real workplaces, and observation can subtly change behavior.
Cost/speed profile: high effort per session, but each session is dense with discovery.
Choose contextual inquiry before you design, when scoping a new product or entering an unfamiliar domain. Choose a usability test after you design, when you have something concrete to evaluate. Mature teams do both: inquiry to frame the problem, testing to validate the solution.
What is guerrilla testing good for, and what isn't it?
Guerrilla testing is a cheap, fast smoke test for early concepts, not a substitute for structured research.
You take a phone or laptop to a café, politely intercept strangers, and run five-to-ten-minute sessions for the price of a coffee. In an afternoon you can catch glaring problems: an unfindable button, an incomprehensible label, a value proposition nobody understands.
Strengths: near-zero cost, same-day results, forces designers out of the building, great for killing obviously bad ideas early.
Weaknesses: participants are whoever happened to be nearby, almost never your actual target users; sessions are shallow; no context; findings can mislead if treated as representative. Testing a radiology tool on café patrons tells you about café patrons.
Cost/speed profile: lowest cost, fastest to start.
Choose it when: you are at the sketch or early-prototype stage, your product serves a broad consumer audience, and you need directional signal before investing in proper recruitment. Never let it be the only research a product gets.
How many users do you need for a usability test?
About five users per round, per distinct audience segment. Then fix what you found and test again.
This is the well-known heuristic from Jakob Nielsen and Tom Landauer's research: roughly five participants surface the majority of usability problems in a given design, and each additional user yields diminishing returns. The often-cited figure is that five users uncover around 85% of issues detectable with that test setup.
The nuance most articles skip:
- Five per segment, not five total. If doctors and patients use the product differently, that is two rounds of five.
- Iterate. The power of the heuristic is in repetition: test five, fix, test five again. Three rounds of five beats one round of fifteen.
- Quantitative studies need more. If you are measuring task success rates or comparing designs statistically, you need 20–40+ participants for meaningful confidence intervals. The five-user guideline applies to qualitative problem-finding only.
- Complex or high-stakes products (medical, financial, safety-critical) warrant larger and more careful samples.
Moderated or unmoderated: how do you decide?
Moderated when you need to understand; unmoderated when you need to measure. If your top question starts with "why" or "how do users think about…", moderate. If it starts with "how many", "how fast", or "which version", unmoderate. Budgets permitting, pair them: a moderated round to find and explain problems, an unmoderated round to verify the fixes at scale.
How does AI change usability testing?
AI compresses the analysis, not the observation. At Things, AI-augmented workflows handle the mechanical layers of research:
- Transcription of every session, in minutes, in multiple languages. No more choosing between watching the participant and taking notes.
- Synthesis support: clustering observations across sessions, flagging recurring friction points, drafting affinity maps that researchers then verify.
- Highlight reels: automatically locating the moments where participants struggled, so stakeholders watch three minutes of evidence instead of five hours of recordings.
What AI does not do is replace watching real users. A model can summarize a transcript; it cannot notice the pause before a click, the sigh, the workaround a participant does not think worth mentioning. Simulated "AI users" are useful for piloting a protocol, not for validating a product. The observation stays human. The paperwork gets faster.
Which usability testing method should you choose? A decision table
| Situation | Best method | Why |
|---|---|---|
| Early concept, broad consumer audience | Guerrilla | Fast, cheap directional signal |
| Understanding a domain before designing | Contextual inquiry | Reveals real workflows and constraints |
| Evaluating a new flow or prototype | Remote moderated | Depth, flexibility, follow-up questions |
| Benchmarking or comparing variants | Remote unmoderated | Scale, speed, consistent metrics |
| Hardware, eye tracking, confidential work | Lab | Control and equipment |
| Post-launch continuous measurement | Remote unmoderated | Repeatable, low overhead |
| Stakeholders don't believe the problem | Lab or moderated remote | Live observation persuades |
What are the most common usability testing mistakes?
- Testing too late. Usability testing after launch is damage control, not design.
- Leading the witness. "Do you find this easy?" produces polite lies. Ask people to do, then watch.
- Treating guerrilla findings as representative. Directional signal, nothing more.
- One big study instead of many small ones. Fifteen users in one round finds the same problems five would, and then the fixes go untested.
- Testing tasks instead of goals. "Click the export button" tests reading comprehension. "Get this report to your manager" tests the product.
- Skipping synthesis. Sessions nobody analyzes are theater. Findings must become decisions.
How does Things choose testing methods for a project?
We are a studio, not an agency, which means we do not sell a standard package and make your problem fit it. Method selection at Things starts from the research question, the product's maturity, and the risk profile. A fintech onboarding flow gets moderated remote rounds with follow-up unmoderated benchmarks. A warehouse logistics tool starts with contextual inquiry, because the interface's real competitor is a clipboard. An early consumer concept might get a guerrilla afternoon before we commit design hours at all.
Usability testing rarely travels alone: we pair it with user interviews, heuristic evaluation, and data analysis, then synthesize everything into decisions the team can act on. Research before opinion: it is why the work has held up across 50+ products and a Red Dot award, from Istanbul and Elazığ to brands worldwide.
FAQ
Is remote usability testing as good as lab testing? For most software questions, yes. Remote moderated testing delivers comparable insight at lower cost with broader recruitment. Lab testing wins when you need controlled conditions, physical devices, eye tracking, or strict confidentiality.
How much does usability testing cost? Guerrilla rounds cost almost nothing beyond time. Remote unmoderated studies cost platform and incentive fees. Moderated and lab studies add recruitment, facilitation, and analysis. The relevant comparison is against the cost of building the wrong thing.
Can I run usability tests without a finished product? Yes. Clickable prototypes, paper sketches, even competitor products all work. The earlier you test, the cheaper the fixes.
Is five users really enough? For finding qualitative problems in one design for one audience segment, five per iterative round is a well-supported heuristic (Nielsen/Landauer). For statistical measurement, plan for 20+.
Can AI replace usability testing participants? No. AI accelerates transcription, synthesis, and highlight reels, and can help pilot a protocol, but validated insight comes from observing real users.
Test with us
If you are deciding between methods, or deciding whether to test at all, talk to us. Things is a digital design and engineering studio, founded 2018, building for leading brands in Turkey and worldwide. Design. Code. Mastery.
Write to hello@things.ist and tell us what you are building. We will tell you what to test first.