Evaluate AI-moderated interview tools with a shared brief, complete calls, and evidence you can inspect. A feature list cannot tell you whether the moderator follows your methodology or whether the resulting findings survive review. The most useful test includes both interview quality and the work required to produce the final deliverable.
User Intuition provides AI-moderated qualitative research for agencies, consulting firms, and research teams. This guide is published by User Intuition; it is an evaluation framework, not an independent vendor ranking. Apply the same tests to User Intuition and every other platform on your shortlist.
Download the AI interview evaluation checklist (PDF) →
What should you prepare before evaluating platforms?
Prepare a live or recently completed brief, a discussion guide, participant criteria, and a description of the required output. Identify the decision the research will inform. Choose a brief your researchers understand well enough to recognize useful evidence, missed opportunities, and unsupported conclusions.
Separate essential requirements from preferences. If a project requires a specific language, participant source, data agreement, or interview format, failure on that requirement should not disappear inside a high average score. A convenient interface cannot compensate for a method that cannot answer the client’s question.
Use internal rehearsals to check setup and participant experience, then evaluate a small set of suitable participants. Record which conversations are rehearsals and which are research evidence. A colleague acting as a difficult respondent can expose workflow problems, but their invented answers should never enter the client findings.
For an agency evaluation, add the commercial context: who operates the platform, how client materials are handled, what the researcher reviews, and how the output becomes a presentation. The User Intuition agency workflow and sample study show concrete starting points for that assessment.
What are the ten checks that matter?
1. Guide adherence
Check whether complete interviews address the intended learning goals. Distinguish mentioning a topic from producing evidence that answers it. A participant may hear every planned question and still provide too little information to support the decision. Record those gaps instead of treating question completion as research success.
Confirm that your guide can be preserved where consistency matters. Where follow-ups are flexible, inspect whether they remain within the brief. Test changes to wording and review their effect before launching broadly. Keep the guide version alongside the resulting evidence.
2. Useful, neutral probing
Useful probing follows what the participant said and clarifies a concrete experience, reason, tradeoff, or contradiction. Read the exchange before and after an interesting quote. Check whether the moderator supplied an interpretation that the participant merely agreed with, or whether the participant introduced the meaning themselves.
There is no universal number of follow-ups that makes an interview qualitative. Ask whether each probe adds relevant evidence without repeatedly steering or exhausting the participant. Test short answers, uncertainty, and refusals as well as fluent, detailed answers. The right response to some questions is to move on.
3. Modality and participant experience
Choose voice, video, or chat according to what the research needs to capture. Test the actual device, connection, and stimulus flow your participants will use. Observe interruptions, recovery from errors, consent, and how easily someone can stop the session.
Voice and video can provide context beyond the transcript, but inferred emotion should remain open to researcher scrutiny. A polished summary of tone is not a substitute for checking what was said. For usability work, inspect the task and screen-sharing experience rather than assuming every interview platform supports the same workflow.
4. Sample fit and quality
Ask how participants are sourced, screened, and checked for duplicates or unsuitable responses. Test your exact eligibility criteria. Panel size is a reach indicator, not evidence that a particular role or experience is available. Keep the source and exclusions documented in the study record.
User Intuition supports your participants or its 4M participant panel. For client customer research, bringing the client’s sample may be appropriate. For specialty audiences, confirm feasibility and price before launch. Include your own recruitment effort and participant incentives in the comparison when you supply the sample.
5. Analysis and evidence traceability
Pick a finding and follow it back to the relevant interviews. Check whether the supporting quotes mean the same thing in context and whether contradictory cases were retained. Review the difference between participant statements, generated interpretation, and the recommendation written by the researcher.
Inspect denominators when the output reports theme counts. Were participants asked comparable questions? Was the topic raised spontaneously or prompted? A percentage within an interviewed sample should not silently become a population estimate. The agency guide to scaled qualitative research explains that distinction.
6. End-to-end timing
Measure time from the approved brief to a reviewed deliverable. Record setup, recruitment, interviews, processing, researcher review, and client revisions separately. A vendor’s fieldwork example may start after contracting and recruitment; your deadline often starts earlier.
Do not promise a client delivery date from a headline turnaround number alone. Ask what depends on sample availability, language, specialist recruitment, and client approval. During the pilot, record both elapsed time and staff time so you can see where delays and effort arise.
7. Complete project cost
Compare the cost of usable research at the same scope. Include the platform, sample, participant payments, required subscription, review labor, and reporting work. Check minimum commitments, included credits, expiry, replacement rules, and any paid services your study needs.
User Intuition Starter voice interviews cost $30 with your sample or $60 with standard panel recruitment, with no monthly fee. Specialty audiences are quoted separately; incentives you arrange for your own sample are additional. Subscription terms are available on the pricing page. See current pricing.
8. Language and local meaning
Test interviews in the languages you plan to use. Ask a fluent reviewer to inspect the original conversation and transcript, including specialist terms and local references. A translated summary can conceal a misunderstood question or an ambiguous response.
User Intuition supports interviews in 80+ languages. Treat that as a capability to test for your audience and guide. Agree who reviews translated material and how quotations will be presented in the client readout, particularly when a small wording difference changes the interpretation.
9. Data handling and client boundaries
Review the vendor’s current documentation for access controls, retention, deletion, data processing, subprocessors, and model-training policies. Involve the people responsible for your client requirements. Record answers and unresolved items rather than assuming a security badge covers every use case.
For agencies, test how projects and access are separated across clients. Confirm who can see recordings, transcripts, and exports. Plan what happens when a project ends or a researcher leaves. The relevant pass/fail criteria come from the data and agreements involved in the actual brief.
10. Reuse across projects
Ask the platform to retrieve an earlier finding and show its source. Then introduce a new study with a conflicting result and inspect whether that difference remains visible. Reuse is valuable when it preserves context, date, and sample; it is misleading when distinct studies become a single undifferentiated answer.
User Intuition’s Customer Intelligence Hub makes study evidence searchable across projects. Evaluate it using your own workflow and access requirements. For agency work, permission to access one client’s research does not automatically imply permission to reuse it for another client.
What does an annotated User Intuition call show?
The public sample study contains complete phone-purchase and grocery-shopping interviews, plus a presentation from a 43-participant shopper study. The examples below refer to the published phone-purchase transcript. Exchange numbers count each displayed speaker turn from one, including the opening greeting.
| Exchange | What happens in the phone call | Evaluation note |
|---|---|---|
| 5-8 | The moderator asks what triggered the purchase, then asks for the sequence; the participant describes a phone breaking while cooking | A general reason becomes a specific event and context that a researcher can inspect |
| 9-12 | The moderator asks about a specific phone versus getting online quickly, then explores the chosen device | Useful detail follows, but the either/or wording could suggest a frame; test a more open prompt |
| 13-20 | Store choice leads into proximity, urgency, and the staff interaction | Follow-ups connect the purchase goal to concrete experience rather than stopping at a broad satisfaction label |
| 27-28 | The moderator asks what could improve; the participant suggests digital signs | The interview creates an opportunity for criticism after positive comments |
This is a review of the visible conversation, not a controlled validation study. The transcript includes disfluencies and apparent transcription errors. Listen to the recording when wording matters; do not silently polish a quote into a stronger claim. The original full guide is not reproduced here, so this example alone cannot certify complete guide adherence.
The second call supplies a useful challenge to overgeneralization. The phone buyer values staff assistance; the grocery shopper prefers shopping independently. An analyst should preserve that difference rather than merge both interviews into a universal recommendation for more staff contact. Context changes the implication of a superficially similar shopping experience.
The public calls demonstrate inspectable output. They do not establish performance for every audience, the prevalence of these preferences among all shoppers, or agency profitability. Use them to form concrete evaluation questions, then run your own guide with appropriate participants.
How should you record the evaluation decision?
Use an evidence log rather than a score based on impressions. For each criterion, record the requirement, what you observed, the supporting call or artifact, and any unresolved question. Assign a reviewer and a decision: ready, revise and retest, or unsuitable for this brief.
| Decision field | What to record |
|---|---|
| Requirement | The condition the client brief needs the tool to meet |
| Evidence | Complete call, transcript exchange, output, or documented term |
| Result | Meets requirement, needs revision, or does not fit |
| Remaining work | Researcher hours, guide changes, sourcing, or manual review |
| Next step | Owner and action required before fieldwork expands |
Agree the essential requirements before comparing vendors. Review observations independently before discussing them as a team. If reviewers disagree, return to the evidence and identify whether they interpreted the requirement differently. A high total score should never override a failed requirement that makes the project unusable.
After the first paid study, add actual spend, setup time, evidence-review hours, client revisions, and whether a second brief fits the workflow. These observations show whether the platform is becoming useful in practice. A good demonstration is a reason to test; a repeatable delivery process is a reason to expand.
Bring a client brief to the evaluation. Download the checklist and annotated call example, explore User Intuition for agencies, or book a demo → to test your methodology and sample requirements.