By Ajitesh

How Kellogg Designed an AI Oral Exam

How Kellogg Designed an AI Oral Exam

Professor Achal Bassamboo has a line he now uses in the operations group at Kellogg: the premium on the right answer has diminished tremendously. You can get the right answer. The question is whether you can convince other people, and whether you can lead with it.

That is not a slogan about cheating. It is a claim about what a business-school exam is for. Bassamboo chairs Kellogg’s operations department and co-directs the MMM program with Northwestern’s design school. He ran a voice AI final in the Kellogg-HKUST Executive MBA, a cohort spread across time zones, on Tough Tongue AI. Students took it. He graded it. The design is worth writing down.

This post is that write-up. The questions he answered are the ones other faculty already have in their heads: what to test when AI can write the paper, how to design the conversation, and how to deploy it without turning the final into a software incident. I recorded the conversation with him. You can watch it, or keep reading.

If you want the broader case for oral assessment, including the objections that keep most courses from running one, start with oral exams in the age of AI. If you want the mechanics of building an agent, there is a guide for professors. What follows is how one exam was actually designed.

What the exam is for

Bassamboo does not describe a single moment when the written final broke. He describes a path. Paper and pencil worked, because it checked, in real time, whether the student had the idea. COVID moved exams onto Canvas and into people’s homes. Open book and open notes followed, which was honest: that is what is available at work. Then ChatGPT, Claude, and Gemini arrived, and the take-home stopped being a signal of thinking.

How the exam broke
  1. 1
    Paper and pencilChecked understanding in real time.
  2. 2
    COVID: take-home exams on CanvasOpen book, open notes. Honest, because that is what work looks like.
  3. 3
    ChatGPT, Claude, GeminiThe right answer became cheap. The take-home stopped signaling thinking.
  4. →
    The new test: can you explain it, defend it, and convince someone?
The path Bassamboo describes, from paper exams to the question the Kellogg ops group is now arguing about.

The ops lunches at Kellogg have been stuck on the same question since. What should an exam look like in a full-time program, a part-time program, an exec program? His answer is specific. A good exam tests whether you understood the concept, whether you can see the trade-off, and whether you can convince the other person to think. At a business school that claims to make leaders, the last of those is not optional.

The old answer is a viva. In India they call it that, and everyone who has sat through one remembers the temperature of the room. One professor, one student, no draft to hide behind. The problem is logistics. Kellogg sections run 65 to 70 students. Exam week is a week. Grades have to go out. A one-on-one oral does not fit.

He was three weeks from locking a syllabus when a LinkedIn post of ours showed up. We got on a Zoom call. His need was a viva that students could take on their own clock, treated like an interview, or like talking to a board. Ours was a conversation platform. The match was not subtle.

The viva math
A human viva
One professor, one student, face to face. 65 to 70 students per section. About one week of exams, then grades are due.
70 one-on-ones in a week?
A viva with an AI examiner
Each student joins when it suits them.
7 amstudent joins
1 pmstudent joins
11 pmstudent joins
Their own time, any time zone. “Treat it like an interview.”
The format was never the problem. The calendar was.

He did not ban AI for the sitting. He told students they could use whatever they wanted, as long as they treated the session as thinking on their feet. These were executives. Presenting an operations argument in real time is the job. The exam was supposed to look like the job.

A written exam still has a virtue he is careful not to throw away. You can write, reread, scratch it out, put it back. That is report writing, and it still matters. The premium on doing it well has just fallen, because a model can produce the polished page. What a model cannot do in a meeting is pause the room, go consult the model, and come back. The people who will separate are the ones who can challenge the status quo, ask the right question, and offer a perspective without that pause.

Video essays already exist in admissions. This is not that. A video essay is one-way. You get three questions, you deliver three answers. The exam he wanted builds the next question from what the student just said: how did you come up with that, can you take this pro further, there is an inconsistency here. The recording then lets him see not only whether they understood, but whether they could articulate it.

Not a video essay
Video essay
Q1→Q2→Q3
One-way. The same three questions for everyone.
AI oral exam
Q1→your answer→the next question comes from what you said
  • “How did you come up with that?”
  • “Can you explain a little more on this aspect?”
  • “That doesn’t fit with what you said earlier.”
  • “Take that one pro and go deeper.”
The follow-up is the part a static script cannot do, and the part that shows real understanding.

It’s not about the right answer, but the journey that you can go through, you can critique it, you can think on your feet, and you can take me through that understanding that you have developed.

That is also, in his view, what agentic work now requires. You have to prompt well, which means you have to know the nooks. The conversation is the skill.

How to design it

He starts from a topic, not from a tool. A good topic is practically relevant and contains a nugget he calls the secret sauce: the idea he needs to know you have cracked. The exam is built around that nugget. It is coated. If you are paying attention in class and someone says “bottleneck,” you will say “bottleneck.” That is not the test. The test is a picture of a business in which you have to decide that the slowest resource, the idea Goldratt called Herbie, is the idea that applies.

The real world does not arrive labeled “operations” or “marketing.” He cannot yet run a holistic MBA exam that withholds the discipline, though he would like to. Inside one course there are already enough concepts: strategic alignment, bottlenecks, queuing, inventory. The student has to choose.

The shape of the sitting follows from that. The agent presents a situation. The student asks clarifying questions about the process. Only then do more details appear. There are glass walls. It can feel open, like a video game that goes to infinity, but the character cannot leave the compound. Context sits on a slide, because people still need a visual cue. The prompt is somewhat open-ended: what is your viewpoint, and how would you apply what you saw in class. Five modules live inside one running context, so it does not feel like five separate conversations. Difficulty rises, the way a paper-and-pencil exam rises, except the point here is uncovering what was hiding.

The extra thing the model gives you is the second question. The student speaks. The agent looks at the answer and decides what to ask next. That follow-up, built live, is what a static script cannot do.

Hints were the design problem he was most scientific about. If you offer a hint, how do you score the person who needed it against the person who did not? He borrowed the idea from Darrell Duffie, who taught him at Stanford. Duffie ran a market in his exams. You could leave a question blank and take zero, or you could come to him, show where you were stuck, and buy a hint at a price. The exam was supposed to be a learning experience, not only a test. Bassamboo still wants students to walk out having assembled something they already knew into a shape they had not seen.

The agent was allowed to help, with limits. If you asked it to clarify, it rephrased. Repeating the same sentence is not clarifying. If you said you could not hear, it explained. It did not hint forever, because there was a time limit and the exam had to move. Some of that worked. Some of it did not. Language models miss. The goal was that nobody left saying only “I don’t know.”

He has a paper on this with a former PhD student, Vikas, and Sandeep Juneja, his undergraduate advisor. Static exams ask a question and mark it right or wrong. Adaptive exams meet the student: harder if they are right, easier if they are not, so you spend time at the edge of what they know. That logic fits hierarchical material. A business school is more spread out. Someone can be good at inventory and weak at service ops. So he did not make the whole exam adaptive across topics. He adapted inside the concept: nudge toward the point, then score whether they picked it up immediately or only after the nudge.

Two kinds of exams
Static.The same paper for everyone, marked right or wrong.
Adaptive.Meets each student at their level.
difficulty questions → your level right answer → harder wrong answer → easier
Adaptive works when skills stack on each other.Business school skills are spread out, so Bassamboo adapts inside a concept, with hints.
The idea behind the hint design, from his paper with Vikas and Sandeep Juneja.

Scoring stayed with him. The agent gave useful summaries and a map of the conversation. It did not assign the grade. Two reasons, and he is blunt about both. First, he has not specified every conversation that could happen, so trusting a score would be his failure of specification, not the model’s. Second, students are more comfortable knowing he watched the recording, at least this first time. That will change. It has not changed yet.

There is also vocabulary. In operations, bottleneck, slowest resource, and Herbie can all be correct. Bottleneck gets a little more. Herbie counts because that is how the class talked. A generic scorer will not hold those distinctions. If the task were “get the number,” he would have trusted the model more. The task was “explain it so that if I repeated you in class, the room would understand.” That is still a human judgment.

Who grades?
AI exam sessionRecorded and transcribed.
↓
AI prepares the fileA summary, pointers, and the full conversation.
↓
Professor watches and gradesStudents know a human scored them. Vocabulary gets his judgment:
bottleneckslowest resourceHerbieall correct, weighted differently
AI assists. The professor decides, at least for now.
Scoring stayed human for two reasons: student comfort on a first run, and course context a generic scorer would miss.

How to deploy it

The oral exam was not the first AI move in the course. That matters more than it sounds.

Students were already going to use language models, so the faculty started using them too. The old pre-class ritual was a passive Google form: think about the idea we will cover, where have you seen it at work. Ten minutes, everyone primed, class is not a cold start. Those forms became custom chatbots with guardrails, questions that go deeper, and, in one of his favorite exercises, a drawing prompt. Students sketch a process as a picture. The bot asks them to make it a flowchart: boxes for activities, triangles for buffers. Input, output, steps, waiting. The learning happens in the exchange, before anyone sits down in the room.

The same idea hit practice problems. The old loop was binary. You got it, or you looked at the solution and killed the problem for life. A chatbot can hint, which keeps the problem alive. The oral exam sat on top of that path. It was not a tool dropped onto a course that had never used AI.

When the format itself is new, he does not assume the environment will work. Kellogg colleagues taught him this when they moved from paper to Canvas. They shipped a fake quiz with no operations content: a clock, a next button, a box to type in. Students learned the room before the grade depended on the room. He does the same before he teaches: he wants to see the classroom.

Oral exams already carry anxiety. A final on a new interface would have mixed two kinds. The fix was a tech check. Students opened a link, gave their name, and talked to the avatar. If the agent addressed them by name, the audio was working. They saw the slides. They learned they needed a quiet room and enough bandwidth. It does not remove exam nerves. It moves them. He wants people slightly anxious. He wants that anxiety to come from the material, not from “what will this tab do when I open it.”

A little anxiety helps
performance anxiety → too relaxedtoo stressed slightly anxious
Nerves should come from the material, not from the tech.
The tech check does not remove exam nerves. It removes the wrong kind.

Two avatars, because people differ. Some would rather talk to the professor. Some would rather not. One voice was his, cloned. The other was a neutral machine voice.

The sitting was split in two, about 25 minutes each. A three-hour viva would have been cruel, and a single short sitting would have cherry-picked one module. He wanted depth across five modules inside one context, and a breather between parts. He gives the credit for chunking to Kellogg colleagues, Marty Lariviere in particular, who break MBA exams into sub-exams so the commitment is not one long sitting.

The avatar’s limits helped. Bassamboo gets excited when he poses a question and leaks the answer. Students know this. The avatar does not slip. The face does not nod. Encouragement is useful in class. In a viva, students hunt for the signal that they have scored the point. The hard vivas of his own student days were the ones where the professor gave nothing away, sometimes even coding the marks so you could see the page and still not know if it was five or one. A less expressive avatar is closer to that examiner than a warm human is. We have also seen, in interviews, that people talk more to an AI than we expected, often looking for an affirmation that never arrives. That changes the dynamics. It is not a reason to avoid the format. It is a reason to watch for it.

He is clear about the downsides. The model has a life of its own. It may not follow the instruction you wrote. Talking to an avatar, even knowing the faculty will watch, is not the same social event as sitting across from the professor. He had not finished the student debrief when we recorded, and his own sample is small. The success claim here is narrow: the exam ran, at scale, for an executive cohort across time zones, and the sessions went smoothly. It is not a randomized trial.

The other gain is time, split rather than saved. He still spent the minutes watching. Students still spent the minutes talking. They did not have to spend them together. Transcript if it is clean, video if it is not. He watched the video.

The shout-out at the end of our conversation was specific. The experiment ran with the KH cohort at HKUST. Judy and Stephanie on the staff there made it possible. Kellogg operations faculty, over years of small conversations, taught him how to write a decent exam. Mesa, Varun’s team, let him mimic what they were already doing in admissions. He had spent a lot of hours with us on the agent. None of that is incidental. An AI oral exam is a faculty design problem that a platform can run. It is not a platform that designs the exam.

What the platform did

Three product choices made the design runnable.

The agent could be built as a conversation with a persona, a question arc, guardrails on how much to reveal, and slides as exhibits. That is the design section above, turned into something a student can open.

It joined Google Meet. Students already knew how to enter a meeting, check a microphone, and speak on camera. An unfamiliar vendor tab would have been another source of anxiety on the day of a final. Meet was the room they already used. Zoom is the same idea. The agent is a participant. The recording, transcript, and analysis come back to the platform.

It had a face. One avatar cloned his voice. One was neutral. The poker face is a feature of that avatar stack, not a side effect of “AI.” For an exam, less expression is useful. For a coaching session, you might want the opposite. The same platform does both. The exam needed the first.

None of that replaces the syllabus decisions. He still chose the nugget, the hints, the two sittings, and the grade. The platform is how those decisions met 65 people who were not in the same time zone.

How the exam ran
  1. 1
    Tech check, before the examSay your name, hear the avatar say it back, see the slides load.
  2. 2
    Join a Google MeetAt whatever time the student picks.
  3. 3
    The examiner joins as an avatarHis cloned voice, or a neutral one.
  4. 4
    The case goes up on a slideThe exhibit stays on screen.
  5. 5
    Answer, follow-up, answerThe next question depends on what the student said.
  6. 6
    Two sittings of about 25 minutesOne context, five modules, a breather in between. Then the professor reviews.
Every step is a decision the professor made. The platform is how those decisions reached a cohort spread across time zones.

Where this goes first

He does not think departments will all converge on one format. Faculty are already trying video walkthroughs of problem solving, locked-down Canvas tests, and oral exams. Edtech will keep investing in testing. Oral exams have a particular leverage: the next question can follow the conversation.

Two things will gate adoption inside a course. The instructor has to spend the time to design the exam. And students have to get comfortable with a grade that came from a conversation a model helped conduct. MCQs are popular because they are black and white. A conversation is not. What did the agent ask, what did it not ask, what did it hint. Until both sides are easy with that, the instructor should own the score. Once they are, he thinks it takes off.

He expects admissions to get there first. A crafted essay is easy to polish. A live conversation, where the question is not even shown until you sit down, is harder to fake as someone else. Some of his students already report a static version: prepare a speech from a prompt, then talk to a machine, then take a follow-up. The better version builds that follow-up from what they said. Recruiting has the same shape: more first-round conversations to get to a shortlist a human then meets. Mesa is already on that path. That is why he wanted to learn from them.

Volume is the other divide. Where the class is small, people will keep doing a human viva or a presentation. Where the class is large and the thing you care about is the thought process, the oral exam has to run through AI or it does not run. Take-home case write-ups, the kind where you send a solution from home, are the format he thinks will shrink first. Controlled paper-and-pencil will not disappear. It tests something else, and it should.

The number 3.75, as an answer, tells him almost nothing. Some people get it for the right reason, some for the wrong one. He wants the reason, and then, even more, the explanation to the next person. The recipe, step one, step two, step three, is not enough. “Do it this way because I went to Kellogg” does not travel. The value is being able to explain how the thing needs to be done.

He landed, late in our conversation, on the PhD defense. The degree is signed off on a document. The examination is still oral. Wherever you want to know that someone has mastered it, nothing beats that. The new part is doing it at the scale of a section, not a committee.

If it was just the steps, then with today’s time you can tell that steps will be done by anyone. The reason is why those steps make sense, because we will rarely get the same situation again. We have to use the concepts from one to the other. That is where the thinking matters much more than just getting the right answer. And that’s the purpose of education.

The main idea

Do not ban the model. Change what you are looking at. Design an unlabeled situation around the secret sauce of the course, let the student ask, let the second question come from their answer, and keep the grade. Put the sitting in a room they already know. Check the microphone before the day. Split the time. Let the face stay still.

If you want a first version you could run this term, the guide to conducting an AI oral exam is the mechanics, and oral exams in the age of AI is the case for doing it at all. This post is the design of one that already ran.

A
Ajitesh
Tough Tongue AI
Share