Chapters

Part 5 · Chapter 18

Evidence of learning

Transcripts turn role play from anecdote into data. Here is what to measure, what to show, and what to distrust.

8 min read · Updated Jul 2026

What you'll learn

  • Growth curves, cohort-level gap analysis, and participation patterns
  • Practicum-hour and PD-compliance logging
  • Why practice count is not learning and how to guard against Goodharting

The artifact problem that has plagued role-play programs since Socrates asked the questions in person: how do you prove it worked? Written exams leave a paper trail. Lab reports leave a notebook. Live role play leaves two memories, and memories disagree.

AI role play leaves everything. Every session produces a recording, a transcript, and a set of rubric scores — the evidence layer that Chapter 4 called out as a fundamental change. That abundance is powerful, but only if you know what to read and what to distrust.

What to measure

Per-student growth curves. The most valuable view is the simplest: a student’s rubric scores across attempts, over time, on the same scenario. Attempt 1 scored 40% on open-ended questioning; attempt 7 scored 82%. That curve is evidence of learning in its purest form — the same task, the same rubric, visible skill acquisition. Show it to the student (it is the most motivating thing in the system), and keep it for the portfolio.

Cohort-level gap analysis. Aggregate rubric scores across the section per criterion and look for the valley. If 80% of the cohort clears “empathic acknowledgment” and 30% clear “concrete next step,” you have found the lesson the course is not teaching. That gap should drive the class debrief (Chapter 13, Layer 4) and, over semesters, reshape the syllabus. This is formative assessment at the program level, something the field has wanted and never had the data to do.

Participation patterns. Who is practicing, when, and how many attempts. A student who ran the scenario twelve times at 2 a.m. the night before the due date has a different practice profile than one who ran it three times across three weeks, and the distinction matters for advising — though not for grading, because the grade is on the best attempt.

Breakthrough rates on specific resistance points. If a scenario has a designed curveball (Chapter 11’s pressure block), track how many students survive it cleanly. This is your scenario’s effectiveness metric: a curveball nobody handles is too hard; one everyone handles is too easy.

What to show your department chair

Three numbers, one semester:

  1. Median growth from first attempt to best attempt, across the cohort. This answers “did students get better?” and is the hardest number for a skeptic to dismiss.
  2. Practice volume per student (mean, median, and range). This answers “did they actually do it?” and demonstrates that the rehearsal gap from Chapter 1 has been addressed for this course.
  3. Rubric-criterion distribution showing where the cohort is strong and where it struggles. This answers “what should we teach differently?” and turns the program review from opinion into evidence.

For accreditation and compliance contexts — teacher PD hours, counseling practicum minutes, clinical simulation logs — session timestamps and durations serve as verifiable records without additional instrumentation. The data your accreditor asks you to generate by hand is generated automatically, which is the administrative case that gets budgets approved even when the pedagogical case is still being debated.

What to distrust

The most dangerous metric is the most obvious one: practice count. It is tempting to require a minimum number of sessions and treat completion as evidence of learning. Resist this. A student who runs the scenario ten times and scores 40% each time has not learned; they have repeated. A student who runs it three times, scores 45%, 62%, 80%, and writes a strong reflection has learned. Count is a proxy for engagement, not for skill. The reflection memo (Chapter 13) exists partly to prevent this Goodharting: requiring a comparison between first and best attempt forces the student to demonstrate growth, not just volume.

A second distrust: cohort averages that mask bimodality. If half the cohort scores 90% and half scores 30%, the mean says “60%, needs work” — but the actual problem is that half the cohort is not doing the practice at all. Check the distribution, not the average, and treat non-engagement as an adoption problem (Chapter 17), not a skill problem.

A third: score inflation over time without scenario refresh. If the same scenario runs for three semesters, the score curve will rise partly because the cohort got better and partly because the scenario leaked (through study guides, group chats, or simply being predictable). Refresh scenarios the same way you refresh exam questions, and track scores per-scenario-version.

On Tough Tongue: instructors see a dashboard with per-student attempt histories, per-criterion score breakdowns, and cohort-level distributions. Session data exports to CSV for compliance reporting or institutional research, and score logs carry timestamps that satisfy PD-hour requirements without a separate attendance system.

Exercise

Define the three numbers you would show your department chair after one semester of AI role play. For each, write the metric, the data source (which rubric criteria, which scenario, which time window), and the claim it supports. Then identify the one metric you are tempted to use but should not — the vanity number that measures activity rather than learning — and write one sentence explaining why it is dangerous. Bring all four to the pilot planning in Chapter 20.