Run a build challenge with dozens of builders and a hard one-week deadline, and by day seven the launches split cleanly into two groups.
The builders who shipped something compelling picked ideas shaped like conversations they already knew how to have. The memoir maker was built by someone who had spent hours interviewing her grandmother. The PM interview coach was built by an actual product manager. The cold-call trainer was built by an ex-SDR who could recite the objections in her sleep.
The builders who stalled picked ideas requiring expertise they didn’t have, or conversations with no natural structure. Their agents sounded plausible for five minutes and fell apart the moment a real user pushed.
The pattern is reliable enough to be a rule: the best first agent is one where you have the domain knowledge, the conversation has a natural structure, and imperfection is acceptable. This chapter turns that rule into a scorecard.
The idea catalog
Three lanes, ordered by observed success rate.
Professional practice has the highest hit rate, because the structure and the content already exist:
- Sales: cold-call practice, objection handling, demo roleplay. Clear rubric, repeatable structure.
- Interviews: PM, behavioral, and case prep. The builder usually has direct experience.
- Customer support: complaint handling, difficult customers. Scripts exist, and the agent’s resistance is the feature, not a bug.
- Training: product knowledge, onboarding, compliance. The content is already documented somewhere in your company.
Creative and personal ideas are tolerant of quirks, with fun failure modes:
- The AI memoir maker: five sessions across life chapters, where memory is the star of the show (Part 3 uses it as a case study).
- The AI dungeon master: persistent characters and adventures, a memory and image-generation showcase.
- Debate partners, language tutors, podcast guest simulators, first-date conversation coaches.
Products to sell are harder, because they add a second customer (the buyer) on top of the user:
- Exam-season micro-coaches for schools, therapy pre-screeners for clinics, ADHD accountability buddies, elderly-care companions with family reports, wedding-toast practice sold through event managers.
- Defer these unless you already have the distribution channel. The agent is the easy half.
Five characteristics of a good first idea
Score your idea against each. Be honest; the transcript will be.
1. Repeatable structure. The conversation follows a shape you can write down: greet, present, practice, feedback, wrap. If every conversation is wildly different, you cannot engineer the context for it. Part 2 will show you exactly why structure is what prompts are made of.
2. You have the domain knowledge. This is the make-or-break one. The domain knowledge is the hardest part of agent building: gaining all the context and engineering it into the agent. Deep-research tools help you map unfamiliar terrain, but do not build in a domain where you have just no clue. Chapter 23 goes deep on research methods; none of them fully substitutes for having been in the room.
3. Tolerance for imperfection. Practice and roleplay tolerate an occasional odd phrasing; a slightly clumsy objection from an AI buyer is still useful practice. Medical advice does not tolerate odd phrasing. Match the stakes of your use case to the current reliability of the technology, and revisit as the technology moves.
4. A clear success signal. You can tell from one transcript whether the session worked: the learner practiced aloud, the objection got handled, the story got captured. If you cannot define success in a sentence, you cannot iterate toward it, and Part 7’s whole improvement loop depends on this sentence existing.
5. Reachable users. You personally know five people who would use it this week. Not a market thesis. Five names.
Scoring: five out of five, build now. Four, build. Three, sharpen the idea before writing a single prompt. Two or fewer, pick another idea; the weeks you save are real.
Scoping v1
Once the idea passes, the discipline flips: build less than you think.
One scenario, one flow, one persona. The multi-scenario course, the difficulty ladder, the persona library: all of it comes later, and Part 4’s sub-agent architecture makes it natural when it does. Version one is a single conversation done convincingly.
Write the success criterion before the prompt. One sentence: “By session end, the user has practiced X aloud twice and received specific feedback.” This sentence becomes your evaluation rubric, your iteration target, and eventually your production quality signal. It is the cheapest artifact in this playbook and the most reused.
Timebox aggressively. With an experience platform, the first working conversation should take under an hour. The first good conversation arrives after roughly ten iterations; Part 6 covers that loop. If you are three days in without a working conversation, the problem is scope, not skill.
Example: one idea through the scorecard. “Cold-call practice for SDRs”: repeatable structure, yes (opener, objections, close). Domain expertise, yes (the builder was an SDR). Tolerance for imperfection, yes (it’s roleplay). Clear success signal, yes (did they handle the brush-off?). Reachable users, yes (her old team). Five for five: build. Notice how fast that took to evaluate. The scorecard is an afternoon filter, not a business plan.
You have an idea and a stack. Now the skill that makes or breaks the agent, the one the next six chapters teach: engineering its context.