Run this test on any role-play scenario before assigning it: play it badly, on purpose. Give vague answers, skip the empathy, ask leading questions, bulldoze. If the persona warms up anyway — if the “angry parent” is mollified by a non-apology, if the “resistant client” opens up to a closed question — the scenario is not just weak. It is actively harmful, because it just rewarded the exact behavior you are trying to train out.
And here is the part that surprises faculty: nothing is malfunctioning. Language models are trained to be helpful, agreeable, and accommodating. Left to its defaults, every AI persona is a pushover, because being a pushover is the model’s job description. Whenever you need it to be difficult the way humans are difficult, you have to do extra work. That work has three ingredients.
Ingredient 1: Resistance
The formula: whenever the student offers an explanation, proposal, or reassurance, the persona pushes back two to three layers deep. It never accepts the first answer at face value.
The canonical example is native to education. A student tells the teacher, “I didn’t do the homework because there was no electricity.” A real teacher doesn’t nod sympathetically and move on. She asks when the power went out. Whether it was out all evening. What the student did instead. Two or three follow-ups deep, either the truth emerges or the effort does.
Written as prompt material, adaptable to any persona:
## RESISTANCE BEHAVIOR
- Never accept the first explanation or reassurance at face value
- Probe 2-3 layers: ask a specific follow-up, then a harder one
- Only soften when the student's response meets the breakthrough
criteria below
- If the student is vague, name the vagueness:
"You keep saying you'll 'support' my daughter. What will you
actually do differently on Monday?"
Only the content changes across personas; the pattern is constant. The angry parent probes for specifics and commitments. The examiner probes reasoning: “you said the model overfit — how did you know?” The reluctant counseling client deflects twice before anything true comes out. The moot-court judge asks where in the statute. Fill the content from your incident file (Chapter 10); the hardest real conversation you ever had in this role is the specification.
Ingredient 2: Breakthrough conditions
Resistance without a win condition is a wall, not a workout. A real difficult counterpart has a teachable moment inside them: I am not accepting what you’re saying, but if you give me enough of the right thing, I can be moved. Define that “right thing” explicitly, or the persona will either never concede or concede at random — and both failure modes teach nothing.
## BREAKTHROUGH CONDITIONS
Begin to soften ONLY when the student:
- Acknowledges my specific concern in their own words (not a
scripted empathy phrase), AND
- Offers a concrete next step with a named time or action
When it happens, show it visibly: my tone drops, I sit back,
I ask a collaborative question ("okay — what would that look
like for the reading assignments?").
Three design notes, because this small block carries the pedagogy:
- The breakthrough criteria are your rubric, enacted. The student wins by performing the observable behaviors from Chapter 10, step 2. This is the deep integrity property: gaming the scenario and mastering the skill become the same act.
- Make the concession visible and rewarding. The moment the hostile parent exhales and says “okay, tell me more” is the payoff that brings students back for attempt six. Don’t let the persona concede in a mumble.
- Gradients beat cliffs. A partial thaw at layer one (“fine, I’m listening”) and full breakthrough deeper in gives students a slope to climb instead of a wall to bounce off.
Unwinnable personas demoralize; pushovers flatter. Everything your scenario teaches lives between those two failure modes, in this block.
Ingredient 3: Pressure
Real counterparts don’t just resist — they derail. The parent’s phone rings. The client suddenly asks, “do you even have kids?” The examinee’s counterpart flips it: “how would you have scoped it?” Scripted personas must therefore go off-script on purpose. Script the categories and the frequency, never the lines:
## PRESSURE (1-2 times per session, at natural moments)
Choose from:
- Personal deflection: "How long have you even been teaching?"
- Emotional spike: brief tears or a flash of anger about an
earlier incident with the school
- Authority test: "I want to speak to the principal about this."
- Scope jump: raises an unrelated grievance mid-conversation
After the curveball, return to the main thread regardless of how
they handled it, and note the handling for the evaluator.
Sample the categories from real incidents, not brainstorms. Invented curveballs have a tell; the ones from your incident file land.
Calibrating the ladder
You now hold three knobs: resistance depth, breakthrough strictness, pressure frequency. Turn them together and you get the difficulty ladder that Chapters 5 and 8 kept promising — beginner, intermediate, and advanced variants of the same persona, three configurations of one role card.
Calibrate with transcripts, not intuition. Two failure smells: students winning in one exchange means resistance is too shallow — deepen it. Students looping on the same answer or going silent means it’s too hard — soften the deepest layer, or add a hint behavior (“if the student stalls twice, offer an opening in character: ‘look, all I really want to know is…’”). For graded scenarios, calibrate on volunteers first; Chapter 17’s self-test exists for exactly this.
One placement rule: resistance belongs to the counterpart persona only. If your session has a warm-up briefing or a feedback phase, those stay kind. Friction is a phase property, not a session property — the angry parent is angry; the coach who debriefs the attempt afterward is not.
Exercise
Take the role card from Chapter 10 and add the three blocks: resistance behavior with one named-vagueness line in your persona’s voice, breakthrough conditions that restate your two most important rubric behaviors, and a pressure menu of three curveballs drawn from real incidents. Then run the bad-student test from this chapter’s opening on the result — play it badly and confirm the persona holds. If you can bulldoze your own scenario, your students certainly will.