Chapters

Part 5 · Chapter 20

A rollout blueprint

One course, one moment, one semester. Then expand.

9 min read · Updated Jul 2026

What you'll learn

  • A semester-by-semester rollout from pilot to program
  • The six failure modes that kill pilots before they prove anything
  • The full playbook arc in one page

The worst thing that can happen to AI role play at a university is a big launch. A dean announces it across five programs simultaneously, faculty receive a platform login and a suggestion to “explore use cases,” students encounter it ungraded and unframed in week three, and by midterms it is furniture — technically available, functionally dead. This is how most educational technology pilots end: not with a verdict, but with a shrug.

The rollout blueprint that works is the one that is small enough to succeed, long enough to learn, and documented well enough to expand. One course, one moment, one semester.

Semester 1: The pilot (one course, one moment)

Weeks 1–4: Design and self-test. Pick the highest-pain, lowest-controversy conversational moment in the course (Chapter 17). Build the scenario with a role card, resistance, and a rubric (Chapters 10–12). Run the self-test: you as student, twice — competent and incompetent — grading your own transcripts. Adjust until the persona holds and the rubric scores fairly. This is four to six hours of design work spread across a few weeks.

Weeks 3–4: Student orientation. Rung 1 of the stakes ladder (Chapter 15): one short, ungraded, explicitly framed session. The syllabus paragraph from Chapter 16 is distributed. The goal is that every student has met the format before anything counts.

Weeks 5–10: Formative use. Rung 2: sessions count for completion, not rubric scores. Students practice at their own pace. The instructor monitors transcripts lightly — not to grade, but to calibrate: is the persona too easy, too hard, or hitting the wrong thing? Adjust difficulty and rubric if needed. Run one class debrief (Chapter 13, Layer 4) around week 7 using anonymized transcript excerpts.

Weeks 11–15: One graded touchpoint. Rung 3: the best-attempt submission with the comparative reflection memo. Or, if the course runs an AI viva, weeks 11–12 for the dry run with volunteers (Chapter 19), weeks 13–14 for the live assessment. Grade with human review over automated scores.

Week 16: Measure and document. Collect the three numbers from Chapter 18 (median growth, practice volume, rubric-criterion distribution). Write a one-page summary for your department: what you did, what you measured, what students said, what you would change. That page is the rollout artifact. Without it, the pilot evaporates; with it, the next semester has a foundation.

Semester 2: Expand within the course

Add a second scenario targeting a different moment. Move the first scenario to rung 2 earlier (it is familiar now) and introduce the new one at rung 1. If the course runs a simulation or an actor session, wire the pre-event AI sparring pattern from Chapter 7 so the live event improves visibly.

Invite one colleague to observe: not to adopt, but to see. The strongest evangelism is a faculty member watching another’s cohort arrive at the actor session having already survived the scenario ten times.

Semester 3: Expand to the program

With two semesters of data, expand to two or three additional courses in the same program. The design principles (Chapters 10–13) are now documented; the first adopter mentors the next ones. Standardize what should be standard (rubric format, LMS embedding pattern, disclosure language) and customize what should be custom (scenarios, resistance calibration, grading weight).

This is also when the program-level gap analysis from Chapter 18 becomes possible: comparing rubric-criterion distributions across courses reveals where the program is strong and where students consistently struggle, which is the accreditation-relevant finding.

Common failure modes

Six patterns kill pilots before they prove anything. Name them so you can avoid them:

1. Launching graded before formative. The stakes ladder exists for a reason. Students who encounter an unfamiliar format at grade weight blame the format, not their skill, and the complaints reach the dean before the data reaches the chair.

2. Skipping the debrief. The most common silent failure. Practice volume climbs, scores plateau, and the program concludes that “AI role play doesn’t work.” In reality, the program ran Stage 2 without Stage 3 (Chapter 2’s skeleton). Mandate the debrief layer before evaluating the program.

3. The agreeable persona. A persona that folds at the first student response teaches nothing and is worse than no practice, because it builds false confidence. The self-test (Chapter 17) catches this if you run it honestly. If veteran faculty can breeze through the scenario, students certainly will.

4. No faculty champion. Technology deployed by an IT department or teaching center without a faculty partner who owns the pedagogy dies of neglect. The champion does not need to be an enthusiast; they need to be a practitioner who cares about the course and is willing to do the design work.

5. Treating it as a content library. Assigning a generic “practice interviews” scenario without tying it to the course’s learning objectives, rubric, and debrief structure is educational technology at its worst: tool without pedagogy. The scenario is not content to consume; it is an exercise to design, assign, debrief, and iterate on.

6. Measuring the wrong thing. Practice count and session count are engagement metrics, not learning metrics (Chapter 18). A pilot that reports “students completed 400 sessions” without reporting growth, rubric scores, or student reflection has measured activity and claimed learning. The three-number framework from Chapter 18 prevents this.

The arc in one page

This playbook argued the following, chapter by chapter:

Students get almost no repetitions at the conversations that matter most (Chapter 1). Every role play has four stages — prepare, perform, debrief, assess — and the performance stage has been the bottleneck for fifty years (Chapter 2). Four modalities can staff the counterpart; each wins in a different context, and the answer is one modality per moment, not per program (Chapter 3). AI changes cost, leakability, and evidence; it does not change the debrief or the design, which stay human (Chapter 4).

Career readiness, clinical training, simulations, language learning, and oral exams are the five use-case families (Chapters 5–9). Scenarios are designed backwards from learning objectives (Chapter 10), made hard through engineered resistance (Chapter 11), scored with concept rubrics (Chapter 12), and consolidated through structured debrief (Chapter 13).

Interaction shape varies from solo drills to projected avatars to real phone calls (Chapter 14). Syllabus integration follows a stakes ladder and four assignment patterns (Chapter 15). Safety, ethics, and equity are design problems with a net-positive equity case (Chapter 16). Faculty resistance is legitimate and addressed through the self-test and the batting-cage framing (Chapter 17).

Evidence comes from growth curves and cohort gap analysis, not session counts (Chapter 18). AI-conducted exams need operational infrastructure, not just a prompt (Chapter 19). And the rollout is one course, one moment, one semester — documented well enough to expand (this chapter).

The rehearsal gap is fixable. Not by replacing what works, but by making what was scarce abundant, and concentrating human expertise where it always should have been: on design and debrief. The professor’s job does not shrink. It relocates.

Exercise

Write the one-page pilot plan for your course. Five sections: (1) the course and the moment, (2) the scenario you built in Chapters 10–13, (3) the three syllabus touchpoints from Chapter 15, (4) the three metrics from Chapter 18, and (5) the colleague you will invite to observe in semester 2. Send it to yourself as a calendar item for the start of next semester. A plan that exists only in this playbook does not count; one that exists on your calendar does.