Chapter 3 decided who plays the counterpart. This chapter decides the physical shape of the encounter: one student with headphones, a whole class facing a screen, a phone ringing at a scheduled time, or an agent sitting in a video call. Faculty coming from actor programs tend to assume one shape — the scheduled, room-based session — because that is the only shape actors come in. AI comes in four, and the choice is per-activity, not per-program.
Shape 1: Solo voice
The default and the workhorse: one student, one persona, a browser tab or phone app, no audience. This is where the volume lives — the interview drills, the language ladders, the practice clients, the pre-viva rehearsals. Its virtues are the ones Chapter 3 scored highest: zero scheduling, zero social cost, unlimited repetition, automatic transcripts.
Choose solo voice whenever repetition is the mechanism. Its one real limitation is worth naming: nobody is watching, which means nobody sees body language, and for skills where posture and eye contact are the curriculum (patient positioning, courtroom presence), solo voice trains the words but not the body. Video-based sessions narrow this; live observation still owns it.
Shape 2: The projected persona
This is the shape faculty most often miss, and the situation flagged throughout this playbook where an avatar beats 1:1 delivery outright: when the shared experience is the point, put one persona on the classroom screen and let the room work it together.
The class collectively cross-examines the expert witness. The seminar interrogates the historical figure they’ve been reading. The cohort of trainee teachers takes turns de-escalating the same projected parent, watching each other’s attempts and the persona’s different responses. The town-hall stakeholder from Chapter 7 takes questions from thirty citizens.
Three properties make the projected shape work. First, it restores the observational learning that solo practice loses — students watch each other try, which every classic guide lists as a core benefit of classroom role play. Second, it creates common material for the Chapter 13 class debrief: everyone saw the same conversation. Third, it is the natural home for an avatar rather than voice alone: a face on a screen gives the room a shared focal point, and the visual persona carries presence that disembodied audio does not. This is also the shape where human-in-the-loop avatar platforms remain genuinely competitive — a live specialist reading the room can improvise off thirty faces in a way AI cannot yet. Use AI-driven for cost and repeatability; splurge on human-driven for the one marquee session a semester where nuance is everything.
A facilitation note: the projected shape needs a human conductor. Someone calls on speakers, pauses the persona for a sidebar (“what should she try next?”), and cuts the session when the moment has landed. The instructor’s Stage 2 role, extinct in solo practice, returns here in full.
Shape 3: The real channel
Sometimes the delivery channel is part of the fidelity. A crisis-line trainee should practice on an actual ringing phone, because answering a phone cold — no video, no name, no context — is the skill. An admissions counselor training on inquiry calls should hear a call, not click a widget. A student preparing for panel interviews should join a video meeting, wait in the lobby, and manage the etiquette, because the meeting mechanics are part of the anxiety being trained away.
The rule: when the scenario’s real-world version arrives through a channel, and the channel contributes to the difficulty, deliver the practice through that channel. When the channel is incidental (most drills), the browser is fine and frictionless wins.
On Tough Tongue: all four shapes run from the same scenario definition. Solo practice through a link or an LMS iframe embed, the projected shape by running the session on the room’s display, phone delivery over a real call for channel-fidelity drills, and a Google Meet agent that joins video sessions for meeting-shaped practice. Designing once and choosing the shape per assignment is the intended workflow.
Shape 4: The embedded step
The quietest shape: the role play living inside something else. A module in the LMS where the reading is followed by a five-minute embedded conversation before the quiz unlocks. A pre-lab safety dialogue. A checkpoint inside an online course where the student must talk through their project plan before proceeding. Here the role play is not an event at all — it is a step in a sequence, and its virtue is that nobody schedules anything. Chapter 15 builds the syllabus patterns around this.
Choosing, quickly
- Repetition is the mechanism → solo voice.
- Shared experience or observational learning is the mechanism → projected persona, avatar on the screen.
- The channel contributes to the difficulty → real channel (phone or meeting).
- The role play gates or follows other coursework → embedded step.
- Body language is the curriculum → video shapes at minimum; keep live observation in the loop.
Mixed designs are normal and usually best: solo drills all term, one projected session mid-term for the shared debrief, the graded attempt through the real channel.
Exercise
Return to the three activities you scored in Chapter 3’s exercise and assign each a shape from this chapter, with one sentence of justification tied to the mechanism (“repetition,” “shared experience,” “channel fidelity,” or “sequence”). Then check the calendar implication: the solo and embedded shapes need no room and no slot, the projected and channel shapes need exactly one each. Most faculty discover their plan needs one scheduled hour a term more than they currently use — and about twenty fewer than the actor version required.