Building a Skill Interview with AI A developer built a voice-based AI skill assessment app in which an AI agent interviews a participant about a chosen skill and produces a scored report card. Skills are defined entirely in JSON files (including agent name, tone, voice, model effort, duration and per-point scoring requirements), so new skills can be added without code changes, and the OpenAI and Azure Speech components sit behind interfaces so providers can be swapped. The developer reports that calibrating the prompt, particularly ending the interview at the right time, was the hardest part. Here's the idea: build an app that checks if someone knows their stuff using an interview. You pick a skill, type your name, and off you go. An AI agent chats with you by voice, asks questions, and at the end you get a report card. Hopefully with more compliments than trauma. I've seen some applications that evaluate interview participants with AI, so I decided to build one out of curiosity. It was an interesting journey. The trickiest part was calibrating the prompt, especially to end the interview at the right time. Let's get the requirements first, so we have a clear picture of what to implement, before I start the architecture, the code, drawing boxes and arrows, and pretending I know what I'm doing lol. Every skill is a JSON file. Adding a new skill means adding a file: no code changes. A skill defines: | Field | Purpose | |---|---| | Title | Name shown to the participant, e.g. "SQL" | | InterviewInstruction | What the agent should ask and which topics matter most | | AgentName , AgentTone | Who the agent is and how it behaves friendly, professional, like a Jedi master... | | AgentVoiceGender , AgentVoiceType | The agent's voice Male/Female; HighPitched, Neutral, Deep, Warm | | Effort | How capable the AI model should be: Low, Medium or High | | ReportInstruction | How to evaluate and what to put in the report | | MaxPoints | The highest possible score | | InterviewDurationInMinutes , ExtraInterviewDurationInMinutes | Planned duration and extra time | | PointInstructionMap | A label and the minimum requirements for every point from 1 to MaxPoints | For example, a Star Wars skill can use fun level names: { "Id": "star-wars", "Title": "Star Wars Lore", "InterviewInstruction": "Test the participant's knowledge of the Star Wars universe...", "AgentTone": "Warm and playful, like a wise old Jedi master.", "AgentName": "Master Oren", "AgentVoiceGender": "Male", "AgentVoiceType": "Deep", "Effort": "Low", "ReportInstruction": "Evaluate breadth and depth of lore knowledge.", "MaxPoints": 4, "InterviewDurationInMinutes": 10, "ExtraInterviewDurationInMinutes": 2, "PointInstructionMap": { "1": { "Label": "Youngling", "Requirements": "Knows the main characters." }, "2": { "Label": "Padawan", "Requirements": "Knows the plot of the main films." }, "3": { "Label": "Jedi Knight", "Requirements": "Knows the history of the Jedi and the Sith." }, "4": { "Label": "Jedi Master", "Requirements": "Knows deep lore, including series and books." } } } If a skill file is invalid for example, a missing point in PointInstructionMap , the app starts with a clear message saying what to fix. This is a sample for an article, so some things are deliberately left out: The interview logic is agnostic and doesn't depend on a specific AI vendor. The interviewer and the report agent OpenAI and the speech-to-text/text-to-speech service Azure Speech are each behind an interface, in their own project, so that they can be replaced by different providers. Let's build the abstractions first: the backend components, and how the interview engine handles everything the participant says so the conversation flows as naturally as possible. The engine is split into two parts: The participant's voice is transcribed and sent to the agent; then the agent writes its reply, and the reply is synthesized into speech with the voice from the config. Every box in the bottom row is an interface, implemented in its own project SkillsValidator.Engine.OpenAI , SkillsValidator.Engine.AzureSpeech . The engine never references a vendor SDK. Skills are read-only: they come from the JSON files. public interface ISkillConfigRepository { Task