AI has changed the economics of UX research. Work that used to take a week (transcribing interviews, coding transcripts, clustering themes) now takes an afternoon. But speed is not insight, and a hallucinated theme is worse than no theme at all.

At Things, our position is simple: research before opinion. We listen, test, and learn. Then we decide. AI belongs in that loop as an accelerant, never as a substitute for real users. This guide covers where AI genuinely helps across the research lifecycle, where it quietly fails, and how to keep your evidence trustworthy while moving fast.

Where does AI fit in the UX research lifecycle?

AI is useful at almost every stage of a research project: planning, recruitment screening, discussion guides, transcription, synthesis, and reporting. The pattern that works: AI drafts, structures, and compresses; humans decide, verify, and interpret. The moment AI starts generating your evidence rather than processing it, you have left research and entered fiction.

Here is the honest map, stage by stage.

Planning. AI is a strong sparring partner for research design. Feed it your product context and business question, and ask it to propose research questions, challenge your assumptions, and flag what your study cannot answer. It is particularly good at stress-testing scope: "Which of these questions actually require moderated sessions, and which could a survey answer?" You still own the design. AI just makes the first draft cheap.

Recruitment and screening. Use AI to draft screener surveys, then to triage open-ended screener responses at scale, flagging professional survey-takers, contradictory answers, and low-effort responses. This is tedious pattern-matching work, and models do it well. Final participant selection stays human, because your sample defines everything downstream.

Discussion guides. AI drafts a competent interview guide in minutes. More valuable: ask it to critique your guide for leading questions, double-barreled questions, and jargon. Then pilot the guide (more on simulation below) before you spend a real participant on discovering that question four makes no sense.

Transcription. Solved. Modern speech-to-text handles most sessions well, including accented English and multilingual conversations. Always spot-check names, product terms, and numbers, because transcription errors compound in synthesis. Never upload recordings to tools you have not vetted for privacy (see below).

Synthesis. The biggest win, and the biggest risk. Covered in depth next.

Reporting. AI turns a validated set of themes into a readable report, an executive summary, and stakeholder-specific cuts (a version for engineering, a version for leadership) fast. Keep the traceability: every claim in the report should link back to real evidence a human has verified.

How do you analyze interview transcripts with AI?

Work in small, verifiable steps: summarize one transcript at a time, extract claims with verbatim quotes attached, then cluster themes across sessions, checking quotes against the source at every step. Never paste ten transcripts into a prompt and ask "what are the themes?" You will get plausible, confident, partially invented output.

The failure mode of AI synthesis is not obvious wrongness. It is plausibility. A model asked to summarize research will produce clean themes with realistic-sounding quotes, some of which don't exist in your data. If you can't trace a theme to a timestamp, it isn't a finding.

A workflow that holds up:

  1. Per-transcript pass. For each interview, ask for a structured summary: participant context, key observations, pains, workarounds, notable quotes (each quote verbatim, with a location reference). Spot-check the quotes against the transcript. This is your quality gate.
  2. Claim extraction. Convert observations into atomic claims ("P3 abandoned checkout because the shipping estimate appeared too late"), each tagged to a participant.
  3. Cross-session clustering. Now let AI cluster claims into candidate themes. Because every claim carries provenance, you can audit any theme back to raw evidence in seconds.
  4. Human interpretation. AI tells you what was said and how often. It cannot tell you what matters; that requires product context, business context, and judgment. Prioritization stays human.
  5. Contradiction hunt. Ask the model to argue against your themes: "Which evidence contradicts this? Which participants are exceptions?" Models are surprisingly good devil's advocates when explicitly asked, and dangerously agreeable when not.

The same pipeline works for open-ended survey data. For hundreds or thousands of free-text responses, AI-assisted coding (classify each response against a codebook you define, with a sample manually verified) is one of the highest-leverage uses of AI in research today. It replaces days of manual coding with hours, while keeping every code auditable.

One rule above all: AI processes evidence; it does not generate it.

Should you use synthetic users and AI-simulated interviews?

Use synthetic users to rehearse your research, never to replace it. Simulated participants are excellent for piloting discussion guides, training junior moderators, and pressure-testing usability protocols before real sessions. As a source of findings about actual humans, they are unreliable, and can be dangerously convincing.

The case for simulation is practical. Real participants are expensive and slow to schedule. Burning your first real session on a broken protocol is waste. Running a simulated session (an AI playing a plausible persona working through your tasks and questions) surfaces confusing task wording, ordering problems, and missing probes cheaply. We use simulated dry-runs at Things for exactly this: rehearsal, not research.

Simulations are also useful for hypothesis generation. A simulated persona might raise an objection you hadn't considered: a question worth asking real users, not an answer.

The case against substitution is structural. A language model produces the statistically plausible response, an average of what people like your persona tend to say in its training data. Real research value lives in the opposite place: the surprising workaround, the emotional spike, the thing nobody predicted. Synthetic users are systematically biased toward the expected. They cannot be surprised by your product, because they have never used it. They mirror the assumptions baked into your persona prompt, so "validating" a design with synthetic users often means agreeing with yourself through a very elaborate mirror.

If a team tells you they replaced user interviews with synthetic users and their findings all confirmed the roadmap, that is not a success story. That is the failure mode, working as designed.

The line: synthetic users test your instruments. Real users test your product.

What are the bias, hallucination, and privacy risks of AI in UX research?

Three risks dominate: hallucinated evidence, amplified bias, and mishandled participant data. All three are manageable with process, and all three will quietly corrupt your research if ignored.

Hallucination. Models invent quotes, merge participants, and smooth contradictions into false consensus. Mitigation: verbatim quotes with references, per-transcript processing, spot-check audits, and a standing rule that unverifiable claims get cut from reports. Treat AI synthesis output as a draft from a fast, overconfident junior researcher: useful, never final.

Bias. AI adds its own layers on top of your existing sampling and confirmation biases. Models trained predominantly on English-language, Western, online text will misread participants outside that distribution. They also tend toward sycophancy: ask "do users find this confusing?" and the framing itself tilts the answer. Use neutral prompts, ask for disconfirming evidence explicitly, and keep humans from underrepresented-in-training-data user groups in the interpretation loop.

Privacy and consent. Interview recordings are personal data. Before any AI touches them:

  • Update your consent forms. Participants should know AI tools will process their recordings, and be able to opt out.
  • Vet your tools. Understand where data goes, whether it trains models, and how long it is retained. Enterprise agreements with no-training clauses or self-hosted models are the safe defaults.
  • Anonymize before processing. Strip names, employers, and identifying details from transcripts before synthesis. It also reduces model bias toward participant demographics.
  • Mind regulation. GDPR, KVKK, and similar frameworks apply. "The vendor probably handles it" is not a compliance posture.

None of this is exotic. It is the same rigor good researchers always applied, extended to a new class of tool.

How does Things pair AI speed with real-user rigor?

We run AI through the entire pipeline (planning, screening, transcription, coding, clustering, reporting), but every finding that reaches a client traces to a real human being, verified by a researcher. That is the whole model: AI compresses the mechanical work; the saved time gets reinvested in more sessions, deeper analysis, and better decisions.

Things is a studio, not an agency: founded in 2018, based in Istanbul and Elazığ, with 50+ shipped products and a Red Dot award on the shelf. "Design. Code. Mastery." isn't decoration; it describes a team that researches, designs, and engineers under one roof. Our research practice covers user interviews, usability testing in every format that fits the question (remote, lab, contextual, guerrilla), heuristic evaluation, and data analysis and insight synthesis.

Concretely, AI-augmented research at Things looks like this:

  • Before fieldwork: AI-drafted guides, critiqued and piloted through simulated sessions before a single real participant is scheduled.
  • During fieldwork: automated transcription and structured session notes, so moderators stay present with the participant instead of typing.
  • After fieldwork: claim-level coding with full provenance, AI-assisted clustering, human interpretation, and reports where every insight links to real evidence.

The result is not "cheaper research." It is more research per decision: faster loops, larger samples, and findings a team can actually trust enough to act on. Research before opinion, at the speed modern product teams need.

FAQ

Can AI replace user interviews? No. AI can prepare, transcribe, and analyze interviews, and simulated sessions can pilot your protocol. But findings must come from real users; synthetic participants reproduce expected answers and miss the surprises that make research worth doing.

What is the best way to analyze interview transcripts with AI? Process one transcript at a time, extract claims with verbatim quotes and references, verify a sample against the source, then cluster claims across sessions. Avoid single-prompt "summarize everything" synthesis; it invites hallucinated themes.

Are synthetic users ever worth using? Yes: for piloting discussion guides, rehearsing usability tests, training moderators, and generating hypotheses to test with real people. They are a rehearsal tool, not an evidence source.

Is it safe to upload interview recordings to AI tools? Only with informed participant consent, a vetted tool that does not train on your data, and anonymization before processing. Enterprise or self-hosted deployments are the safer default for sensitive studies.

How much time does AI actually save in UX research? It depends on the study, but transcription, coding, and first-draft synthesis (historically the slowest phases) compress from days to hours. The rigor-preserving move is to reinvest that time in more sessions and deeper verification, not to ship faster on thinner evidence.

Ready to research faster, without cutting corners?

If you want AI-augmented research with real-user rigor behind it (interviews, usability testing, synthesis, and the design and engineering to act on what you learn), talk to us.

Things, a studio, not an agency. Design. Code. Mastery. Write to hello@things.ist.