AI-Powered Conversational Evaluation
Co-Creating with Machines: AI in the Creative Industries
A full-day event exploring the intersection of artificial intelligence and creative practice. The day featured six speakers, interactive breakout sessions using Padlet, and culminated in a live performance combining improvised music, AI-cloned voice narration, real-time visuals built in Unreal Engine, and audience-generated poetry. The performance was designed to demonstrate how AI can function as a collaborative creative partner rather than a replacement for human expression.
Simplified for Everyone
The DIY method presented below has been substantially simplified from the original project, which required advanced technical skills in game engines, real-time audio analysis, and custom software development. The step-by-step guide strips this back to the core principle, using AI to have natural conversations with participants, so that anyone with access to a phone or laptop can experiment with using AI to help evaluate engagement activities.
AI-Powered Conversational Evaluation
Replace rigid post-event questionnaires with a natural, adaptive conversation led by a large language model, letting participants share what actually mattered to them.
Phase 0: Logic Model & Topic Mapping
Identify the “Evaluation Gap” this method fills. It is best suited for experiential or arts-based events where feelings are hard to reduce to a Likert scale and small organisations with limited evaluation capacity, time, or trained interviewers.
Primary Objective
To surface genuine attitudes, emotional responses, and unanticipated themes from participants through natural dialogue, rather than pre-determined questions that constrain what people can tell you.
Key Stakeholders
Project leads who need evaluation evidence but lack a trained interviewer. Participants who find forms off-putting or inaccessible. Funders who want rich qualitative data alongside numbers.
Reflective Planning Questions – Ask Yourself Before You Begin:
- “What are the three to five things I most need to understand about how participants experienced this, and can I express them in plain language for the AI?”
- “Are my participants comfortable talking to a screen, or would voice mode feel more natural than typing?”
- “Am I willing to let the conversation go in unexpected directions, knowing I might discover something I did not think to ask about?”
Part 1: The DIY Step-by-Step Method
Write the Evaluation Brief
Before your event, craft a system prompt that tells the LLM exactly what it needs to know: what the experience was, who the participants are, and what themes you want to explore. This is where your evaluation design lives. The better your brief, the better the conversation.
The Details:
- Action: Open any LLM you have access to (ChatGPT, Gemini, Claude, or a free alternative). Write a prompt that describes the event, lists the evaluation themes you care about, and instructs the AI to have a natural, flowing conversation rather than firing questions one after another.
- Tip: Tell the AI explicitly to let the participant steer the conversation at times. The richest data comes when people go off-script. Include a line like: “Allow the participant to go in unexpected directions. Follow their lead before gently returning to the evaluation themes.”
Set Up the Conversation Station
Create a simple, low-friction setup where participants can walk up and start talking. This could be a tablet on a table, a laptop in a quiet corner, or a phone propped up with a sign. The less it looks like a feedback booth, the better.
The Details:
- Action: Set up a device with your pre-loaded LLM prompt. If using voice mode (available on paid tiers of most LLMs), make sure the microphone works and background noise is manageable. If using text, increase the font size so it feels like a chat, not a form.
- Tip: Place a short, friendly instruction card next to the device. Keep it to one or two sentences: “Tell this AI what you thought of today. Just chat like you would with a friend.” Avoid words like “evaluate,” “feedback,” or “survey” on the card.
- Important: Be upfront that participants are talking to an AI, not a person. Some people are uncomfortable with AI, and surprising them with it will break trust. A simple note on your instruction card saying “This is an AI chatbot” is enough. If someone does not want to use it, respect that and have a low-tech alternative ready, even if it is just a pen and a postcard with an open question on it.
Let the Conversations Flow
Step back and let participants engage at their own pace. The AI will adapt to each person, following their interests while gently probing the themes you care about. Some conversations will be two minutes. Others might run to ten. Both are valuable.
The Details:
- Action: Do not hover. If participants see a facilitator watching them talk to a screen, they will perform for you instead of being honest with the AI. Check in periodically to make sure the device is still working, but otherwise let people come and go freely.
- Tip: If you want to increase participation, offer a creative incentive. For example, tell participants the AI will generate a short poem or image from their conversation at the end, something personal they can take away. This gives them a reason to engage that has nothing to do with “giving feedback.”
Harvest and Synthesise the Data
After the event, collect all conversation transcripts. Feed them back into an LLM and ask it to identify recurring themes, notable quotes, shifts in attitude, and areas of disagreement. This is where the method pays off: dozens of unique conversations distilled into clear patterns.
The Details:
- Action: Copy all conversation logs into a single document. Paste them into an LLM with a new prompt: “Here are transcripts from evaluation conversations with participants at [your event]. Identify the top five themes, note any surprising or contradictory responses, and summarise the overall attitude of participants toward [your topic].”
- Tip: Ask the LLM to flag where its inferences are strong versus speculative. This builds a layer of honesty into your analysis. You can then validate the speculative ones by checking them against other evidence you hold, like observation notes or attendance data.
- Important: Do not send AI-generated analysis straight into a final report. Many funders, ethics boards, and partner organisations are cautious about AI-led interpretation of participant data, and rightly so. LLMs can miss nuance, flatten contradictions, or confidently present something that is not quite right. Treat the AI output as a first draft. A human should read every transcript, cross-check every theme the AI identified, and make the final call on what goes into your report. If a funder asks how the data was analysed, you want to say “AI-assisted, human-verified.”
Part 2: The Realist Pathway
Evidence Manifest
Documenting how the Mechanism was verified through specific data collection.
| Evidence Type | What it Proves | Data Source |
|---|---|---|
| Behavioural Conversation Engagement Rate |
Participants willingly engage with the AI without facilitator prompting, demonstrating that the conversational format lowers barriers to participation compared to traditional feedback forms. | Count of completed conversations versus total attendees. Comparison with response rates from conventional post-event surveys at similar events. |
| Qualitative Thematic Depth and Range |
Conversations surface themes and emotional responses that do not appear in structured surveys, proving the adaptive format reaches areas a fixed questionnaire cannot. | AI-generated thematic analysis of all conversation transcripts, cross-referenced with any parallel survey data to identify themes unique to the conversational method. |
| Attitudinal AI-Inferred Attitude Summaries |
The LLM can reliably infer participant attitudes toward key themes from conversational data, validated by asking participants to confirm or correct the summary at the end of their conversation. | End-of-conversation attitude summaries generated by the LLM, paired with a simple participant confirmation (“Does this sound right to you?”) to measure inference accuracy. |
Don’t just write a boring report. Use the stuff you actually made during the session.
The Theme Cloud
Feed all your transcripts into an LLM and ask it to rank the most frequent themes by how often they appeared. Then ask it to generate a simple word cloud or ranked list. This gives funders a single visual that says “here is what your participants actually talked about” rather than “here is what we asked them.”
The Attitude Shift Table
If your LLM prompt asked participants about their feelings at the start and end of the conversation, you can build a simple before-and-after table showing how attitudes shifted during the exchange. Even three or four of these create a compelling narrative of change.
The Unexpected Finding
The biggest strength of this method is that participants tell you things you did not think to ask. Pull out two or three insights that no survey would have captured and present them as “discoveries.” Funders love evidence that an evaluation method found something new, not just confirmed what was already assumed.
The Voice of the Room
Ask your LLM to generate a short creative piece from the conversation transcripts, a poem, a collective statement, or a “letter from participants.” This is not fluff. It is a powerful reporting artefact that captures tone and feeling in a way a bar chart never will. Drop it into your report as a full-page pull quote or use it to open a presentation.
The “One-Sentence” Report
“Because we used AI-Powered Conversational Evaluation, participants in our community felt heard and curious enough to share honest reflections they would never have written on a form.”