AI-Powered Conversational Evaluation

AI-Powered Conversational Evaluation – Evaluation Method 7
About This Project

Co-Creating with Machines: AI in the Creative Industries

A full-day event exploring the intersection of artificial intelligence and creative practice. The day featured six speakers, interactive breakout sessions using Padlet, and culminated in a live performance combining improvised music, AI-cloned voice narration, real-time visuals built in Unreal Engine, and audience-generated poetry. The performance was designed to demonstrate how AI can function as a collaborative creative partner rather than a replacement for human expression.

Location Queen’s University Belfast, Media Lab and OMA Buzz
Participants AI practitioners, creative industry professionals, researchers, and students (approx. 50 attendees)
Delivered by Daniel Brice, QUB Media Lab, in collaboration with performers Franziska Schroeder and Ailish Brice (IDEA Lab / SiCi)
A Note on This Method

Simplified for Everyone

The DIY method presented below has been substantially simplified from the original project, which required advanced technical skills in game engines, real-time audio analysis, and custom software development. The step-by-step guide strips this back to the core principle, using AI to have natural conversations with participants, so that anyone with access to a phone or laptop can experiment with using AI to help evaluate engagement activities.

Evaluation Method: 7

AI-Powered Conversational Evaluation

Replace rigid post-event questionnaires with a natural, adaptive conversation led by a large language model, letting participants share what actually mattered to them.

Phase 0: Logic Model & Topic Mapping

Identify the “Evaluation Gap” this method fills. It is best suited for experiential or arts-based events where feelings are hard to reduce to a Likert scale and small organisations with limited evaluation capacity, time, or trained interviewers.

Primary Objective

To surface genuine attitudes, emotional responses, and unanticipated themes from participants through natural dialogue, rather than pre-determined questions that constrain what people can tell you.

Key Stakeholders

Project leads who need evaluation evidence but lack a trained interviewer. Participants who find forms off-putting or inaccessible. Funders who want rich qualitative data alongside numbers.

Reflective Planning Questions – Ask Yourself Before You Begin:
  • “What are the three to five things I most need to understand about how participants experienced this, and can I express them in plain language for the AI?”
  • “Are my participants comfortable talking to a screen, or would voice mode feel more natural than typing?”
  • “Am I willing to let the conversation go in unexpected directions, knowing I might discover something I did not think to ask about?”

Part 1: The DIY Step-by-Step Method

1

Write the Evaluation Brief

Before your event, craft a system prompt that tells the LLM exactly what it needs to know: what the experience was, who the participants are, and what themes you want to explore. This is where your evaluation design lives. The better your brief, the better the conversation.

The Details:
  • Action: Open any LLM you have access to (ChatGPT, Gemini, Claude, or a free alternative). Write a prompt that describes the event, lists the evaluation themes you care about, and instructs the AI to have a natural, flowing conversation rather than firing questions one after another.
  • Tip: Tell the AI explicitly to let the participant steer the conversation at times. The richest data comes when people go off-script. Include a line like: “Allow the participant to go in unexpected directions. Follow their lead before gently returning to the evaluation themes.”
Sample System Prompt: “You are helping to evaluate a creative event. You will have a relaxed, natural conversation with someone who just attended. Your goal is to understand how they felt during the experience, what moments stood out, and whether anything changed in how they think about [your topic]. Do not use a questionnaire format. Instead, chat naturally, ask follow-up questions based on what they say, and let them lead at times. At the end, summarise your sense of their overall attitude and the key themes that emerged.”
2

Set Up the Conversation Station

Create a simple, low-friction setup where participants can walk up and start talking. This could be a tablet on a table, a laptop in a quiet corner, or a phone propped up with a sign. The less it looks like a feedback booth, the better.

The Details:
  • Action: Set up a device with your pre-loaded LLM prompt. If using voice mode (available on paid tiers of most LLMs), make sure the microphone works and background noise is manageable. If using text, increase the font size so it feels like a chat, not a form.
  • Tip: Place a short, friendly instruction card next to the device. Keep it to one or two sentences: “Tell this AI what you thought of today. Just chat like you would with a friend.” Avoid words like “evaluate,” “feedback,” or “survey” on the card.
  • Important: Be upfront that participants are talking to an AI, not a person. Some people are uncomfortable with AI, and surprising them with it will break trust. A simple note on your instruction card saying “This is an AI chatbot” is enough. If someone does not want to use it, respect that and have a low-tech alternative ready, even if it is just a pen and a postcard with an open question on it.
3

Let the Conversations Flow

Step back and let participants engage at their own pace. The AI will adapt to each person, following their interests while gently probing the themes you care about. Some conversations will be two minutes. Others might run to ten. Both are valuable.

The Details:
  • Action: Do not hover. If participants see a facilitator watching them talk to a screen, they will perform for you instead of being honest with the AI. Check in periodically to make sure the device is still working, but otherwise let people come and go freely.
  • Tip: If you want to increase participation, offer a creative incentive. For example, tell participants the AI will generate a short poem or image from their conversation at the end, something personal they can take away. This gives them a reason to engage that has nothing to do with “giving feedback.”
4

Harvest and Synthesise the Data

After the event, collect all conversation transcripts. Feed them back into an LLM and ask it to identify recurring themes, notable quotes, shifts in attitude, and areas of disagreement. This is where the method pays off: dozens of unique conversations distilled into clear patterns.

The Details:
  • Action: Copy all conversation logs into a single document. Paste them into an LLM with a new prompt: “Here are transcripts from evaluation conversations with participants at [your event]. Identify the top five themes, note any surprising or contradictory responses, and summarise the overall attitude of participants toward [your topic].”
  • Tip: Ask the LLM to flag where its inferences are strong versus speculative. This builds a layer of honesty into your analysis. You can then validate the speculative ones by checking them against other evidence you hold, like observation notes or attendance data.
  • Important: Do not send AI-generated analysis straight into a final report. Many funders, ethics boards, and partner organisations are cautious about AI-led interpretation of participant data, and rightly so. LLMs can miss nuance, flatten contradictions, or confidently present something that is not quite right. Treat the AI output as a first draft. A human should read every transcript, cross-check every theme the AI identified, and make the final call on what goes into your report. If a funder asks how the data was analysed, you want to say “AI-assisted, human-verified.”

Part 2: The Realist Pathway

Inputs
Ingredients: One device (phone, tablet, or laptop) with access to any LLM. A written evaluation brief describing the event and themes. A short instruction card for participants. A quiet or semi-private space near the event.
Mechanism
The Spark: People disclose more in conversation than on forms because dialogue feels reciprocal, not extractive. A natural back-and-forth lowers self-consciousness and lets participants explore what mattered to them rather than what you assumed would matter. The AI adapts in real time, following threads a paper survey could never anticipate, while gently returning to your evaluation themes. This creates a hyper-personalised experience where every participant feels heard, not processed.
Outcomes
The Impact: Rich, thematic qualitative data that captures genuine attitudes and emotional responses. Unexpected insights that fall outside your original evaluation framework. A dataset that can be rapidly synthesised by the same AI tools, turning hours of analysis into minutes. And participants who leave feeling they had a meaningful exchange, not that they filled in a form.

Evidence Manifest

Documenting how the Mechanism was verified through specific data collection.

Evidence Type What it Proves Data Source
Behavioural
Conversation Engagement Rate
Participants willingly engage with the AI without facilitator prompting, demonstrating that the conversational format lowers barriers to participation compared to traditional feedback forms. Count of completed conversations versus total attendees. Comparison with response rates from conventional post-event surveys at similar events.
Qualitative
Thematic Depth and Range
Conversations surface themes and emotional responses that do not appear in structured surveys, proving the adaptive format reaches areas a fixed questionnaire cannot. AI-generated thematic analysis of all conversation transcripts, cross-referenced with any parallel survey data to identify themes unique to the conversational method.
Attitudinal
AI-Inferred Attitude Summaries
The LLM can reliably infer participant attitudes toward key themes from conversational data, validated by asking participants to confirm or correct the summary at the end of their conversation. End-of-conversation attitude summaries generated by the LLM, paired with a simple participant confirmation (“Does this sound right to you?”) to measure inference accuracy.
Reporting Hacks: Turning Conversations into Proof

Don’t just write a boring report. Use the stuff you actually made during the session.

The Theme Cloud

Feed all your transcripts into an LLM and ask it to rank the most frequent themes by how often they appeared. Then ask it to generate a simple word cloud or ranked list. This gives funders a single visual that says “here is what your participants actually talked about” rather than “here is what we asked them.”

THE HACK: Drop the ranked theme list into a free word cloud generator and screenshot it for the front page of your report.
The Attitude Shift Table

If your LLM prompt asked participants about their feelings at the start and end of the conversation, you can build a simple before-and-after table showing how attitudes shifted during the exchange. Even three or four of these create a compelling narrative of change.

THE HACK: Present each shift as a single row: “Participant arrived feeling [X], left feeling [Y], because [Z emerged in conversation].”
The Unexpected Finding

The biggest strength of this method is that participants tell you things you did not think to ask. Pull out two or three insights that no survey would have captured and present them as “discoveries.” Funders love evidence that an evaluation method found something new, not just confirmed what was already assumed.

THE HACK: Frame each unexpected finding as: “We did not ask about [X], but [number] participants raised it independently.”
The Voice of the Room

Ask your LLM to generate a short creative piece from the conversation transcripts, a poem, a collective statement, or a “letter from participants.” This is not fluff. It is a powerful reporting artefact that captures tone and feeling in a way a bar chart never will. Drop it into your report as a full-page pull quote or use it to open a presentation.

THE HACK: Paste all transcripts into an LLM and prompt: “Write a short poem or collective statement that captures the mood, language, and key feelings expressed across these conversations.”

The “One-Sentence” Report

“Because we used AI-Powered Conversational Evaluation, participants in our community felt heard and curious enough to share honest reflections they would never have written on a form.”

Your Voice Matters|Built by Practitioners|For Practitioners

Help us improve this tool.

Tell us who you are, what you used, and what would make it better. It takes two minutes.

Give Feedback →