By Ajitesh

How to Conduct an AI Oral Exam: A Guide for Professors

How to Conduct an AI Oral Exam: A Guide for Professors

In this AI world, student assignments are becoming harder to interpret as evidence of understanding. A policy paper may mean careful thinking, or it may mean heavy AI assistance. Which one is it? In our conversations with professors, this has become a genuinely difficult part of teaching, and often the only way to find out is to ask the student, “Why did you make this choice?”

That question does not have to be written. It has to be asked. How do you do that at scale? That is where oral assessment becomes an opportunity in higher education. Oral exams have centuries of history in universities and classrooms, and because of AI, the format is getting a fresh look. I just came across an article from Cornell that describes how oral exams, assignment discussions, mock interviews, oral defenses, and role-play scenarios can be a better way to assess learning. Its guidance also recommends pairing spoken assessment with papers, problem sets, or projects when instructors want students to explain their work directly. Cornell’s oral assessment guide is a useful sign of how broad the category has become.

Voice AI now lets a professor conduct these five-minute conversations personally, at scale. It can speak in the professor’s voice with their avatar, which feels hyper-realistic. It can engage with the student, explain the assignment, probe assumptions, play the role of a patient or a client, and capture the exchange for review and analysis.

Byond the automated final exam, the more interesting use case is a series of short conversations across the semester. A student can practice privately, get feedback, try again, and demonstrate what they know, and as a professor you also get to see how they are improving.

We have been working with some leading universities, including the Kellogg School of Management, where Tough Tongue AI was used to conduct a voice AI exam, and SMU Dallas, where it is being used as part of an assignment. We have seen strong results on learning outcomes.

If you are still weighing whether oral assessment is worth the time, I have written a broader piece on oral exams in the age of AI that covers the evidence, the objections faculty actually raise, and where the format breaks down. This guide assumes that decision is made and focuses on the mechanics.

This guide explains the design system behind how professors can use AI in the oral exam context.

What is an AI oral exam?

For the purpose of this blog, an AI oral exam is a spoken assessment conducted with a voice agent. You, as the professor, define the learning objective, question flow, and rubric based on source material and what was taught in class.

The student answers aloud. The agent listens, asks probing follow-up questions, and at the end you get a recording, a transcript, and evidence for review.

The format can be used in several ways:

  • An oral assignment defense after a paper, project, presentation, or code submission
  • A short knowledge check on the week’s material
  • A viva that asks the student to explain and extend an argument
  • A professional simulation with an AI patient, client, parent, employee, or customer
  • A structured reflection after an internship, lab, placement, or group project

These activities do not all carry the same risk. A five-minute practice conversation that gives private feedback is different from an AI-led final worth 40 percent of a grade. The design and oversight should change with the stakes.

Start with weekly practice

The most common mistake is to introduce voice AI at the most consequential moment in the course. Students meet an unfamiliar interface, an unfamiliar examiner, and an unfamiliar assessment format at the same time. Even a technically good system will feel unfair.

A 2026 NYU study shows both the promise and the warning. Panos Ipeirotis and Konstantinos Rizakos used voice AI to conduct 36 personalized oral exams for an undergraduate AI and machine learning course. The total model cost was $15, or $0.42 per student. Seventy percent of students agreed that the format tested genuine understanding. At the same time, 83 percent found it more stressful than a written exam, and 83 percent had never taken any oral exam before. The agent also bundled questions, failed to randomize reliably, and used a cloned voice that some students experienced as aggressive. The full paper is worth reading because it publishes the failures as well as the results.

The practical response is to make spoken interaction part of the weekly rhythm before using it for assessment. Familiarity lowers avoidable stress. Repetition also turns the system into a learning tool instead of a verification tool that appears only when the professor is suspicious.

A simple progression looks like this:

  1. Give every student one ungraded orientation session.
  2. Assign short weekly practice conversations for completion.
  3. Discuss anonymized examples in class.
  4. Let students repeat the scenario and compare attempts.
  5. Introduce one graded oral defense after the format is familiar.

This is the same low-stakes-first pattern I recommend in our rollout blueprint for AI role-play. One course, one conversational moment, one semester is enough for a serious first pilot.

Four voice AI activities that fit a course

1. Oral assignment defense

The student submits a normal artifact, such as a paper, spreadsheet, slide deck, design, lab report, or code repository. The voice agent receives that artifact as context and asks questions about specific choices.

Consider a computer science assignment. A weak oral defense asks, “Explain your code.” A stronger version asks why the student chose one data structure, what would break at ten times the input size, and which part they would rewrite after seeing the result. The second version tests understanding of the student’s own work rather than their ability to recite a summary.

This is a strong first use case because it sits beside an assignment the course already has. The professor does not need to redesign the whole syllabus.

2. Weekly oral knowledge check

The student has a five-minute conversation about the week’s concepts. The agent begins with recall, then asks for an example, a comparison, or an application to a new situation.

For an economics course, the first question might ask the student to define price elasticity. The follow-up should not ask for a longer definition. It could describe a product with few substitutes and ask the student to predict how demand would respond to a price change. The shift from recall to application is where the conversation becomes useful.

These checks should usually be formative. Let the student retry, and score completion or reflection rather than the first raw performance.

3. AI role-play for professional practice

In this format, the AI is a counterpart rather than an examiner. A nursing student speaks with a worried patient. A teacher candidate handles a parent conference. A business student negotiates with a supplier. A social work student conducts an intake conversation.

The scenario needs resistance. A parent who immediately agrees with the teacher does not test the student’s ability to listen and de-escalate. A good persona has a hidden concern, pushes back in a consistent way, and changes only when the student demonstrates the target behavior.

UNSW has described an AI conversation simulator for patient interactions, parent-teacher interviews, and language learning. Its scenarios can change the avatar’s state as the conversation develops, which shows how discipline-specific practice is moving beyond static chat. The EDUCAUSE project description also emphasizes that the activities are repeatable and available outside class.

4. Structured oral reflection

The agent asks the student to explain what happened during a project, placement, simulation, or internship. It probes decisions, contributions, surprises, and what the student would change.

This works better than “tell me what you learned” when the agent has a structure. It can ask the student to describe one difficult moment, name the assumption they made at the time, and connect the experience to a concept from the course.

Reflection also closes the practice loop. A conversation without debrief can simply repeat a bad habit. Our guide to designing the debrief explains why self-assessment should come before automated feedback.

How to design an AI oral exam step by step

Step 1: Define the evidence you need

Start with the learning objective, then translate it into something a student can say or do in a conversation.

“Understand stakeholder management” is too vague. “Can identify an unstated concern, check their interpretation, and adjust a recommendation” is observable. A question flow and a rubric can be built around that behavior.

The scenario design method we use has four parts: objective, observable behavior, forcing situation, and role card. It keeps the technology downstream of the teaching decision.

Step 2: Choose one spoken format

Do not ask the first activity to be an exam, role-play, presentation, reflection, and interview at once. Pick one.

For a first pilot, I would use a five to ten-minute oral assignment defense. It has a clear source artifact, a narrow purpose, and a transcript the professor can review quickly.

Step 3: Build a predictable question arc

A useful oral assessment usually needs five phases:

  1. Orientation: explain the format, recording, time, and stop conditions.
  2. Opening: ask the student to state the core idea in their own words.
  3. Probe: question one choice, claim, assumption, or piece of evidence.
  4. Transfer: change one condition and ask the student to adapt.
  5. Reflection: ask what they would reconsider or improve.

Publish this structure in advance. The exact follow-up questions can vary, but the type of thinking should not be a surprise.

Step 4: Ground the agent in course material

Give the agent the syllabus section, assignment instructions, learning objectives, rubric, and any approved reference material. For an oral defense, also give it the student’s submission.

Keep those sources separate from the behavior instructions. “Use this rubric” and “ask one question at a time” are different kinds of rules. The NYU study found that prompt instructions alone did not reliably enforce one-question turns or random selection. Important constraints should be handled by the workflow or code where possible.

Step 5: Write a concept-based rubric

Avoid criteria that reward magic words. “Mentions informed consent” is easy to game and may miss a student who handles consent correctly in different language. A stronger criterion asks whether the student explained the choice, checked understanding, and respected the relevant boundary.

For each criterion, define what strong, partial, and insufficient evidence sounds like. Require the system to cite moments from the transcript instead of returning a score with no explanation. Our guide to concept-based grading goes deeper into this distinction.

Step 6: Test with strong and weak answers

Take the exam yourself twice. In the first run, answer like a prepared student. In the second, give vague, polished answers that avoid the substance.

The agent should respond differently. It should move forward when there is enough evidence, probe when an answer is incomplete, and avoid rescuing the student by supplying the answer. If both runs feel the same, the assessment is not ready.

Step 7: Give students practice and a clear disclosure

Students should know:

  • That the counterpart is AI
  • What is recorded and transcribed
  • Who can review the session
  • How long the data is retained
  • Whether the activity is formative or graded
  • How to request an accommodation or alternative format
  • How to stop or report a technical problem

They should also see the format and rubric in advance. UCL’s 2026 guidance on inclusive oral assessment stresses fairness, transparency, consistency, practice, and the risk of anxiety or unintended barriers. The UCL toolkit is worth reviewing before a graded pilot.

Step 8: Keep consequential judgment human

For weekly practice, automated feedback can be immediate and detailed. For a graded defense, the system can prepare a draft score with transcript citations. The professor should review the evidence, adjust the result where needed, and own the final grade.

This distinction matters. A student can tolerate imperfect coaching during a retryable exercise. An unexplained error in a final grade creates an academic decision that needs moderation and appeal.

A sample weekly voice AI plan

Voice AI does not need to appear every week to become part of the course. A small recurring pattern is enough.

WeekActivityTimeStakesEvidence
1Orientation conversation4 minutesUngradedCompletion only
3Concept explanation5 minutesFormativeTranscript and self-reflection
5Professional role-play8 minutesFormativeRubric feedback and retry
7Oral defense of a submitted assignment8 minutesLow stakesProfessor-reviewed evidence
9Harder version of the role-play10 minutesFormativeFirst-attempt and retry comparison
11Final oral defense or performance10 minutesGradedRecording, transcript, rubric, human decision

The class debrief belongs between attempts, not after the grade. Show two anonymized approaches to the same difficult moment and ask students what each person noticed, missed, and changed. The transcript gives the discussion something concrete to examine.

How this can work with Canvas

There are two reasonable levels of Canvas integration.

The basic version uses a normal Canvas assignment. Students open the voice activity through a link or embedded tool, complete the session, and submit a reflection or confirmation. Canvas also supports audio and video media submissions, so a course can run a simpler recorded oral assignment without a live AI agent.

The integrated version uses LTI 1.3. A purpose-built tool can launch from the course, receive the user’s role and course context, connect the activity to an assignment, and return a result to the gradebook. 1EdTech’s LTI overview explains the standard and its role in secure access and grade return.

Canvas itself is moving toward conversational assignments. Instructure and OpenAI announced an LLM-enabled assignment in which educators define learning goals, guide the interaction, capture evidence, and return that evidence to the Gradebook. The announced experience is chat-based, so it should not be described as a voice exam. Still, its structure is relevant: educator-defined objectives, visible evidence, and integration with the existing course workflow. Instructure’s announcement shows where LMS-native conversational work is heading.

There is also a third path that has opened up recently: agent-to-agent integration through MCP. Tough Tongue AI ships an MCP server and agent skills, and Canvas has MCP servers built by third parties. Connect both to the same AI agent, and the whole workflow becomes natural language. You can say, “Create an oral defense scenario for this week’s assignment, then create a Canvas assignment that links to it and publish it to my section,” and the agent builds the scenario on Tough Tongue AI, creates the assignment on Canvas, and wires the two together. There is no custom integration to build. The agent is the integration.

For a pilot, do not wait for a perfect integration. Make the assignment easy to find, keep sign-in simple, and decide where the transcript, rubric, and final grade will live. If the activity proves useful, LTI and grade return become the next step.

Conducting the exam inside Google Meet or Zoom

One underrated design decision is where the conversation happens. Tough Tongue AI agents can be invited into Google Meet and Zoom calls, where they join as a participant the way a human examiner would.

This matters more than it sounds. Students already know how to join a meeting link, check their microphone, and speak on camera. The NYU study found that most students had never taken an oral exam before, and an unfamiliar interface adds stress on top of an unfamiliar format. Running the exam inside a meeting tool the student has used all semester removes one source of surprise, and the infrastructure is reliable because it is the same infrastructure the university already runs classes on.

This is how the Kellogg School of Management conducted its voice AI exam. The agent joined through our Google Meet deployment, and the sessions went smoothly in large part because students were already comfortable speaking and working in Google Meet. The agent conducts the conversation in the meeting, and the recording, transcript, and analysis flow back to the platform for review.

For scheduled exams, the agent can be added to the calendar event so it joins automatically at the right time. For practice sessions, a student or professor can drop in a meeting link and have the agent join on demand.

Accessibility, privacy, and fairness are part of the assessment

Oral assessment can reveal understanding that written work misses. It can also introduce barriers related to speech, hearing, language, anxiety, neurodivergence, processing time, and access to a quiet space.

Separate the construct you want to assess from the communication channel. If the course is testing economic reasoning, an accent or speaking speed should not quietly become part of the score. If the course is testing clinical communication, voice and interaction may be relevant, but the rubric should say exactly how.

Offer the alternatives approved by your institution. These may include additional time, text mode, a human-led session, a different schedule, or another way to demonstrate the same outcome. The University of Delaware’s guide to designing and conducting oral exams also recommends practice opportunities and provides sample questions and rubrics across disciplines.

Treat recordings and transcripts as education records. Limit access, define retention, disclose any model providers, and make the appeal path clear. These decisions belong in the pilot plan, not in the cleanup after a student raises a concern.

What a good first pilot looks like

A good pilot is deliberately narrow. One professor chooses one course and one recurring conversational moment. Students receive an orientation and at least two formative attempts. The professor reviews a sample of transcripts, checks whether the rubric captures the intended learning, and documents what changed across attempts.

The pilot should answer three questions:

  1. Did student performance improve on the target behavior?
  2. Did the conversation reveal evidence the existing assignment missed?
  3. Could the professor review that evidence without creating more work than the activity was worth?

Session count is not enough. Four hundred conversations show usage. They do not prove learning.

The main idea

The useful category is broader than an AI oral exam. It is conversational learning and assessment: students practice, explain, perform, and reflect through dialogue.

Voice AI makes those conversations repeatable. It does not decide what is worth learning, write the final rubric, resolve an accommodation, or own a contested grade. Those remain human responsibilities.

At Tough Tongue AI, we are building for both sides of this loop: repeatable AI role-play for practice and structured oral conversations that give faculty reviewable evidence. The best place to start is still small. Pick one conversation your students need more often, make it safe to practice, and learn from one semester before expanding.

A
Ajitesh
Tough Tongue AI
Share